Skip to content

Summary states lack time/population coverage metadata; equal schemas can't prove a merge or reuse is safe #571

Description

@zzylol

Problem

A #535 schema says what a summary is, not what it summarizes.

#535 gives every ASAP edge one Schema. For a summary edge it records the field layout and the committed state type:

pub struct Schema {
    pub fields: Vec<Field>,            // e.g. job: Plain(Utf8), state: Sketch(KLL{k=200}, PerSubpopulationInstance)
    pub time_index: Option<ColumnId>,  // position of a timestamp column, not a time range
    pub unique_keys: Vec<Vec<ColumnId>>,
    pub closed: bool,
}

Nothing in it says which time range or which population (label values) the state was built from. #535 left this out on purpose ("filters, reduction/group keys, … execution timing, window framework … are not additional Schema or Field members"). Composition is where this becomes a gap: once the planner combines existing summary states (#560 SummaryMerge, reuse of ingested panes, sub-DAG sharing in #537), the state's origin is no longer visible in its producer, and the schema is all that is left. Every example below uses two states with byte-for-byte equal schemas:

Schema(job: Plain(Utf8), state: Sketch(KLL{k=200}, PerSubpopulationInstance)), result_kind = State

Example 1: time. Equal schemas, different answers

Input A Input B Merging A and B is…
latency, [00:00, 00:01) latency, [00:01, 00:02) correct: p99 over [00:00, 00:02)
latency, [00:00, 00:02) latency, [00:01, 00:03) wrong: every observation in [00:01, 00:02) is counted twice, which skews the quantile and doubles counts or frequencies
latency, [00:00, 00:01) latency, [00:02, 00:03) correct only for [0,1) ∪ [2,3); wrong if the result is used for the continuous window [00:00, 00:03)

time_index is a column position. A KLL state has no timestamp column at all, so time_index is None in all three rows, and the schema cannot tell these cases apart.

Example 2: population (label values). Equal schemas, different answers

Input A Input B Merging A and B is…
region='us' region='eu' correct: p99 for us ∪ eu within each job
region='us' tier='premium' wrong: premium US requests are counted in both inputs
region='us' region='us' wrong: everything is counted twice

region is a filter label, not an output column, so it never appears in the schema. The job field only says the state is grouped by job. It does not say which jobs or which rows contributed.

Example 3: time and population together

A = us × [0,1) and B = eu × [1,2). The merged state covers exactly those two blocks. Describing it as {us,eu} × [0,2) (the result of storing a time range and a label set separately) would claim EU data for [0,1) and US data for [1,2) that was never read. The metadata has to keep time and population paired per region.

Example 4: answering a query from a stored state

Query: p99(latency) WHERE region='us' AND ts IN [10:00, 10:05) GROUP BY job. A stored state with the matching schema could hold US data for 10:00–10:05, EU data, or US data for only 10:00–10:03. All three have the same schema. Today the planner can confirm that the state type fits, but not that the contents fit.

Conclusion. Schema equality is necessary but not sufficient for composing or reusing summaries. Without time/population metadata, the planner must either refuse every composition or accept silent double counting and missing data.

What a fix needs

Proposed in #567. Follow-up check of declared populations against subtree filters: #570.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions