Conversation
…ts fields Name the metadata coverage to match #560 and the docs; drop the single-variant multiplicity and deployment-specific revision; rename grouping to reduction to match SummaryAgg; report failures through SchemaDerivationError::Coverage; revert the unrelated PaneCoverageError rename. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…me bounds validate_structure rejects a SummaryAgg without coverage (CoverageError::Missing). CoverageRegion time bounds become optional so tabular sources without a time column can declare coverage. Population stays trusted; #570 tracks checking it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Given required coverage on summary nodes, input and reduction duplicated the producing SummaryAgg fields; drop them along with ProducerMismatch. Type source as Source, matching Scan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SourceCoverage names the rows a physical scan reads for cost comparison, not which observations a summary state holds; rename it so it is not confused with SummaryCoverage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
TODO: should write a design doc for it, not developer doc, human written. |
Move it to docs/design_docs/proposals with problem and motivation, requirements, design, alternatives and key code interfaces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tive schema design Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The design document and the docs it consolidates are reviewed separately on main. This PR keeps code, tests and the ScanSelection rename in docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
According to this PR doc, I still don't think we are solving the right problem here. This is indeed annoying but it is not schema problem. Since the responsibility of schema is mostly to encode data structure, not data semantic. To determine if the above two sums need to be merged, we can either
|
Yes, this is not a schema problem, I will change the PR problem description title. Actually in the code implementation, the "what is summarized" is a field in the Node, in parallel with the field "schema" in the node struct. |

Closes #571.
Problem: a #535 schema says what a summary is, not what it summarizes
#535 gives every ASAP edge one
Schema. For a summary edge it records the field layout and the committed state type:Nothing in it says which time range or which population (label values) the state was built from. #535 left this out on purpose ("filters, reduction/group keys, … execution timing, window framework … are not additional
SchemaorFieldmembers"). Composition is where this becomes a gap: once the planner combines existing summary states (#560SummaryMerge, reuse of ingested panes, sub-DAG sharing in #537), the state's origin is no longer visible in its producer, and the schema is all that is left. Every example below uses two states with byte-for-byte equal schemas:Example 1: time. Equal schemas, different answers
[00:00, 00:01)[00:01, 00:02)[00:00, 00:02)[00:00, 00:02)[00:01, 00:03)[00:01, 00:02)is counted twice, which skews the quantile and doubles counts or frequencies[00:00, 00:01)[00:02, 00:03)[0,1) ∪ [2,3); wrong if the result is used for the continuous window[00:00, 00:03)time_indexis a column position. A KLL state has no timestamp column at all, sotime_indexisNonein all three rows, and the schema cannot tell these cases apart.Example 2: population (label values). Equal schemas, different answers
region='us'region='eu'us ∪ euwithin eachjobregion='us'tier='premium'region='us'region='us'regionis a filter label, not an output column, so it never appears in the schema. Thejobfield only says the state is grouped by job. It does not say which jobs or which rows contributed.Example 3: time and population together
A =
us × [0,1)and B =eu × [1,2). The merged state covers exactly those two blocks. Describing it as{us,eu} × [0,2)(the result of storing a time range and a label set separately) would claim EU data for[0,1)and US data for[1,2)that was never read. The metadata has to keep time and population paired per region.Example 4: answering a query from a stored state
Query:
p99(latency) WHERE region='us' AND ts IN [10:00, 10:05) GROUP BY job. A stored state with the matching schema could hold US data for 10:00–10:05, EU data, or US data for only 10:00–10:03. All three have the same schema. Today the planner can confirm that the state type fits, but not that the contents fit.Conclusion. Schema equality is necessary but not sufficient for composing or reusing summaries. Without time/population metadata, the planner must either refuse every composition or accept silent double counting and missing data.
What this PR adds
Schemastays the layout contract from #535 and does not describe coverage. Coverage is a sibling field on the node, next toschema:coverageis required on summary nodes:validate_structurerejects aSummaryAgg(and, in #560, aSummaryMerge) whose coverage isNonewithCoverageError::Missing. Plain relational nodes leave itNone; the field is anOptiononly because all operators shareOperatorNode.Coverage cannot go inside
Schema. #560 only allows a merge when the input schemas are equal, and the inputs of a useful merge ([0,1)+[1,2)) always have different coverage.Every observation in a region is assumed to contribute once to the state.
Removing duplication
Given the new definition, coverage holds only what no other type records: time × population. This PR also removes the duplication that the first draft introduced or exposed:
inputandreductiondropped from coverage. They copiedSummaryAgg.input/SummaryAgg.reductionon the same node, andwith_coverageneeded aProducerMismatchcheck to keep the copies in sync. feat(ir): define compatible logical summary merges #560'sSummaryMergenow compares them on its producers throughOperatorNode::summary_update().sourceis aSource, not aString. It uses the same type asScan.source, so one table cannot have two spellings.SourceCoveragerenamed toScanSelection(asap-aware-mapping, about 100 call sites plus docs). It names the rows a physical scan reads for cost comparison, a different concept fromSummaryCoverage.revisionandmultiplicitywere removed earlier, as a deployment concern and a single-variant enum respectively.How the examples come out under
merge_disjoint:[0,1)+[1,2), same population[0,2)[0,1)+[2,3)[0,2)+[1,3)PossibleOverlapregion=us+region=eu, same timeregion=us+region=usPossibleOverlapregion=us+tier=premiumPossibleOverlap. Different label names prove nothing.us×[0,1)+eu×[1,2){us,eu}×[0,2)SourceMismatchSummaryMerge(#560)Rules:
validate_structure.Somewithregions = []means known empty.time_ms: Noneis for sources without a time column (plain tabular data). Such a region overlaps every region it is not population-disjoint from.with_coveragevalidates the declaration and requiresStateoutput (CoverageError::NotState).validate_structurere-checks it.with_coverage.Population and time bounds are trusted
Coverage is declared by the composition rule or catalog that built the subtree. Nothing checks
populationagainstSummaryAgg.filter,Filternodes orScan.predicates. Time bounds can't be checked either, becauseTimeRangestores a relativeDuration. So a wrong declaration passes:Follow-up #570 adds the check: the declared population must exactly equal the
column = literalpredicates collected between theSummaryAggand itsScan, and any other predicate shape fails closed. It starts strict about which operators may sit on that path (onlyFilterandTimeRange), becauseProjectorJoincan rename columns or change rows. The list is widened when a real SQL or PromQL plan needs it.Out of scope
Stack and validation
Order: #567 → #560 (
SummaryMergerequires and derives coverage) → #537 (logical transport and CSE preserve them) → #539 → #540 → #561.Tests in
crates/types/tests/summary_coverage.rscover adjacency, gaps, population disjointness, joint regions, overlap, regions without time bounds, source mismatch, serde round-trip, invalid declarations, required coverage, and clearing after rewrites. Each example in this body and the doc is also built as a realSummaryAgg→SummaryMergeplan in #560'ssummary_coverage_examples.rs. The design document is reviewed separately in #573.🤖 Generated with Claude Code