You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A #535 schema says what a summary is, not what it summarizes.
#535 gives every ASAP edge one Schema. For a summary edge it records the field layout and the committed state type:
pubstructSchema{pubfields:Vec<Field>,// e.g. job: Plain(Utf8), state: Sketch(KLL{k=200}, PerSubpopulationInstance)pubtime_index:Option<ColumnId>,// position of a timestamp column, not a time rangepubunique_keys:Vec<Vec<ColumnId>>,pubclosed:bool,}
Nothing in it says which time range or which population (label values) the state was built from. #535 left this out on purpose ("filters, reduction/group keys, … execution timing, window framework … are not additional Schema or Field members"). Composition is where this becomes a gap: once the planner combines existing summary states (#560SummaryMerge, reuse of ingested panes, sub-DAG sharing in #537), the state's origin is no longer visible in its producer, and the schema is all that is left. Every example below uses two states with byte-for-byte equal schemas:
Schema(job: Plain(Utf8), state: Sketch(KLL{k=200}, PerSubpopulationInstance)), result_kind = State
Example 1: time. Equal schemas, different answers
Input A
Input B
Merging A and B is…
latency, [00:00, 00:01)
latency, [00:01, 00:02)
correct: p99 over [00:00, 00:02)
latency, [00:00, 00:02)
latency, [00:01, 00:03)
wrong: every observation in [00:01, 00:02) is counted twice, which skews the quantile and doubles counts or frequencies
latency, [00:00, 00:01)
latency, [00:02, 00:03)
correct only for [0,1) ∪ [2,3); wrong if the result is used for the continuous window [00:00, 00:03)
time_index is a column position. A KLL state has no timestamp column at all, so time_index is None in all three rows, and the schema cannot tell these cases apart.
Example 2: population (label values). Equal schemas, different answers
Input A
Input B
Merging A and B is…
region='us'
region='eu'
correct: p99 for us ∪ eu within each job
region='us'
tier='premium'
wrong: premium US requests are counted in both inputs
region='us'
region='us'
wrong: everything is counted twice
region is a filter label, not an output column, so it never appears in the schema. The job field only says the state is grouped by job. It does not say which jobs or which rows contributed.
Example 3: time and population together
A = us × [0,1) and B = eu × [1,2). The merged state covers exactly those two blocks. Describing it as {us,eu} × [0,2) (the result of storing a time range and a label set separately) would claim EU data for [0,1) and US data for [1,2) that was never read. The metadata has to keep time and population paired per region.
Example 4: answering a query from a stored state
Query: p99(latency) WHERE region='us' AND ts IN [10:00, 10:05) GROUP BY job. A stored state with the matching schema could hold US data for 10:00–10:05, EU data, or US data for only 10:00–10:03. All three have the same schema. Today the planner can confirm that the state type fits, but not that the contents fit.
Conclusion. Schema equality is necessary but not sufficient for composing or reusing summaries. Without time/population metadata, the planner must either refuse every composition or accept silent double counting and missing data.
What a fix needs
Each summary state records which observations it holds: the source, plus time × population, kept paired per block (Example 3).
Problem
A #535 schema says what a summary is, not what it summarizes.
#535 gives every ASAP edge one
Schema. For a summary edge it records the field layout and the committed state type:Nothing in it says which time range or which population (label values) the state was built from. #535 left this out on purpose ("filters, reduction/group keys, … execution timing, window framework … are not additional
SchemaorFieldmembers"). Composition is where this becomes a gap: once the planner combines existing summary states (#560SummaryMerge, reuse of ingested panes, sub-DAG sharing in #537), the state's origin is no longer visible in its producer, and the schema is all that is left. Every example below uses two states with byte-for-byte equal schemas:Example 1: time. Equal schemas, different answers
[00:00, 00:01)[00:01, 00:02)[00:00, 00:02)[00:00, 00:02)[00:01, 00:03)[00:01, 00:02)is counted twice, which skews the quantile and doubles counts or frequencies[00:00, 00:01)[00:02, 00:03)[0,1) ∪ [2,3); wrong if the result is used for the continuous window[00:00, 00:03)time_indexis a column position. A KLL state has no timestamp column at all, sotime_indexisNonein all three rows, and the schema cannot tell these cases apart.Example 2: population (label values). Equal schemas, different answers
region='us'region='eu'us ∪ euwithin eachjobregion='us'tier='premium'region='us'region='us'regionis a filter label, not an output column, so it never appears in the schema. Thejobfield only says the state is grouped by job. It does not say which jobs or which rows contributed.Example 3: time and population together
A =
us × [0,1)and B =eu × [1,2). The merged state covers exactly those two blocks. Describing it as{us,eu} × [0,2)(the result of storing a time range and a label set separately) would claim EU data for[0,1)and US data for[1,2)that was never read. The metadata has to keep time and population paired per region.Example 4: answering a query from a stored state
Query:
p99(latency) WHERE region='us' AND ts IN [10:00, 10:05) GROUP BY job. A stored state with the matching schema could hold US data for 10:00–10:05, EU data, or US data for only 10:00–10:03. All three have the same schema. Today the planner can confirm that the state type fits, but not that the contents fit.Conclusion. Schema equality is necessary but not sufficient for composing or reusing summaries. Without time/population metadata, the planner must either refuse every composition or accept silent double counting and missing data.
What a fix needs
Schema, because a useful merge ([0,1)+[1,2)) has equal schemas but different coverage, and feat(ir): define compatible logical summary merges #560 requires equal input schemas.SummaryMerge, pane reuse, CSE in feat(ir): add logical sub-DAG sharing and export without execution timing #537) can conservatively prove that inputs are disjoint, and rejects them when it can't.Proposed in #567. Follow-up check of declared populations against subtree filters: #570.
🤖 Generated with Claude Code