Skip to content

refactor: unify value and summary-state schemas - #535

Merged
zzylol merged 7 commits into
mainfrom
stack/528-01-schema
Oct 2, 2026
Merged

zzylol merged 7 commits into
mainfrom
stack/528-01-schema

Conversation

@zzylol

@zzylol zzylol commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Part 1/8 of the review stack replacing #528; implements #511's common value/state schema contract (§2.1). Base: main, including #511 and #532.

Before this PR: ordinary operators used Column/Schema, while summary edges used separate SummaryField/SummarySchema/SummaryFamilyType types.

After this PR: both paths use Schema, Field, and FieldDataType. A field can represent a plain value or typed summary state; all existing planner/frontend/runtime consumers use the shared representation. The old operator graphs remain active until later stack layers.

Most changed files mechanically migrate schema fields and constructors. Review the schema definition and raw-source boundary first: SQL qualifiers must survive source binding, as verified by the existing grouped-SUM execution regression. No unified operator or scalar graph is introduced here.

Validation: workspace compilation and tests, the SQL source-binding regression, formatting, and workspace/all-target/all-feature clippy. Subsequent PRs introduce operator/scalar types, graph infrastructure, language lowerings, execution, planner cutover, and legacy cleanup. #528 remains the reference implementation until the stack is complete.

Code interface design: schema, fields, and column references

One edge schema for logical and ASAP-aware plans

Ordinary values and summary state use the same ordered schema. The core declarations in this PR are:

pub type ColumnId = usize;

pub struct Field<T = FieldDataType> {
    pub name: String,
    pub dtype: T,
    pub nullable: bool,
    pub table: Option<String>,
}

pub struct Schema {
    pub fields: Vec<Field>,
    pub time_index: Option<ColumnId>,
    pub unique_keys: Vec<Vec<ColumnId>>,
    pub closed: bool,
}

Field describes one column; it holds no row values or runtime data arrays. The default Field<FieldDataType> is used on operator edges. Nested list and struct members use Field<DataType>, so nested elements cannot carry summary state.

Metadata Meaning
Field.name Producer output name; used for name-based binding.
Field.dtype Plain value type or committed summary-state type.
Field.nullable Whether this field permits NULL.
Field.table Optional SQL table/alias qualifier; preserved for qualified binding across joins.
Schema.fields Ordered field list; its positions define ColumnId.
Schema.time_index Optional position of the time-axis field.
Schema.unique_keys Alternative sets of positions that each uniquely identify a row. An empty outer list makes no uniqueness claim.
Schema.closed Whether the field list completely describes the output. false permits additional dynamic columns, as with schemaless PromQL sources.

An open schema must not be validated as if unlisted labels were absent. Unique-key and time positions must be preserved or remapped when an operator changes the output layout; matching field names alone does not establish those properties.

A field is metadata; a column reference identifies a position

Frontend expressions initially use name-based references:

pub enum ColumnRef {
    Named(String),
    Qualified { table: String, name: String },
    SampleValue,
    Wildcard,
}

Binding resolves these references against a Schema to positional ColumnIds. ColumnId = 2 means the third field in that particular scope, and the corresponding value in an execution row. It is not a globally stable identity across projections or joins and does not prescribe a storage layout. SampleValue denotes the implicit PromQL sample column; Wildcard represents an all-columns/rows request such as COUNT(*).

Schema::column_id(name) returns the first matching position; column_id_qualified(table, name) matches both qualifier and name. Preserving Field.table is therefore necessary to distinguish a.k from b.k.

ASAP summary metadata lives in the field type

The shared type vocabulary is:

pub enum FieldDataType {
    Plain(DataType),
    ExactAggregate(ExactKind, ExactParams),
    Sketch(SketchKind, GroupingStrategy),
    Sample(SamplingKind, SamplingParams),
    Wavelet(WaveletKind, WaveletParams),
    StatModel(StatModelKind, StatModelParams),
}

Plain(DataType) is an ordinary readable value. DataType covers Null, Int64, Float64, Utf8, Bool, timestamps, intervals, dates, lists, structs, and maps. Every other FieldDataType variant describes unfinalized state, preserving its family and concrete configuration:

State family Metadata carried in Field.dtype
ExactAggregate Exact accumulator kind and matching parameter variant: sum, count, min, max, increase, rate, or instant rate.
Sketch SketchKind contains category, concrete algorithm, and algorithm parameters; the field also carries GroupingStrategy. Examples: KLL with k, CMS with width/depth, HLL with precision, or UnivMon with heap size, rows, columns, and layers.
Sample Sampling kind and parameters, such as reservoir capacity.
Wavelet Wavelet kind and parameters, such as Haar retained coefficient count.
StatModel Model kind and parameters, including the deployment-interpreted parametric family name.

SketchKind has private category/algorithm/parameter fields and exposes new(algorithm, params), category(), algorithm(), and params(). Its constructor classifies the category and rejects algorithm/parameter-variant mismatches by assertion. GroupingStrategy distinguishes PerSubpopulationInstance from SharedMultiSubpopulation { kind: HydraKind, params: HydraParams }. It describes the summary layout; actual group-key positions remain operator parameters and ordinary fields in the output schema.

These types provide the identity needed for compatibility checks: KLL and CMS state differ, and two KLL states with different parameters differ. Defining the metadata here does not itself implement summary merging or every declared family in a runtime.

Example: grouped KLL state and its readout

let kll = SketchKind::new(
    SketchAlgorithm::Kll,
    SketchParams::Kll { k: 200 },
);
let state_schema = Schema::lifted(
    vec![
        Field::plain("job", DataType::Utf8, false),
        Field::new(
            "summary",
            FieldDataType::Sketch(
                kll,
                GroupingStrategy::PerSubpopulationInstance,
            ),
            false,
        ),
    ],
    None,
);

This example is closed, has no time-axis field, and makes no unique-key claim because lifted does not infer one. Field 0 is a readable group label; field 1 is typed KLL state. A summary readout exposes a plain result field, such as Plain(Float64) for a quantile. Exact accumulator state similarly requires a finalization boundary before ordinary value consumers use it.

Schema and field metadata describe the edge's shape and state identity. Update expressions, filters, reduction/group keys, requested readout statistic, accuracy guarantees, execution timing, window framework, retention, placement, and cost estimates are separate operator or planning metadata; they are not additional Schema or Field members in #535.

Construction, inspection, and serialization

  • Field::new(...) accepts either vocabulary; Field::plain(...) wraps an ordinary DataType; with_table(...) adds a qualifier.
  • plain_dtype() returns None for state; expect_plain_dtype() panics if a caller incorrectly treats state as a readable value. is_plain() and Schema::is_all_plain() allow explicit checks.
  • Schema::new(fields) starts open with no time axis or unique keys. with_time_index(...) supplies time/key metadata and also starts open. lifted(fields, time_index) starts closed with no unique-key claim.
  • Serialization writes the unified fields representation. Deserialization accepts both legacy columns with bare value types and legacy summary fields with tagged types. Missing closed defaults to open for the legacy columns layout and closed for the fields layout; inputs containing both lists or neither are rejected.

Source at this PR's reviewed head: schema.rs, expr_ir.rs, and sketch.rs.


Stack 1/8 · Previous: main · Next: #536 · Reference/tracker: #528

Rebased onto main at 7734c68f (#544 DAG naming), preserving schema JSON compatibility and the reviewed Field/ColumnId distinction. The #511 design documents now explicitly distinguish metadata from column references: schema contract and expression contract.

Comment thread crates/asap-aware-mapping/src/accuracy/reconciliation.rs
zzylol and others added 5 commits October 2, 2026 19:32
- Schema deserializes the old pre-ASAP layout (`columns`, bare dtypes) and
  the old post-ASAP `SummarySchema` (no `closed`, read as the closed
  `Schema::lifted` shape), so saved plans survive the upgrade.
- dag-viewer render.py reads `fields` (falling back to `columns`); viewer.js
  unwraps `{"Plain": ...}` dtypes for display.
- dag_export: `--table-schema` error names the `columns` key it reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Selvomega
Selvomega previously approved these changes Oct 2, 2026
Comment thread crates/asap-physical-operators/src/physical_planner/precompute.rs Outdated
@zzylol
zzylol merged commit 6e02c7d into main Oct 2, 2026
3 checks passed
@zzylol
zzylol deleted the stack/528-01-schema branch October 2, 2026 20:24
@zzylol

zzylol commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Self-reflection: This PR raises the question of how we should design and name schemas and physical layouts (e.g., tables) for ASAP primitives, or at least summaries.

Calling for a single, complete design document covering schema metadata, physical data, and summary state in #545.

FYI @milindsrivastava1997 @Selvomega

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants