Skip to content

feat(planner): enforce summary response latency bounds - #554

Draft
zzylol wants to merge 2 commits into
stack/528-11-deployment-inputsfrom
stack/528-12-latency
Draft

zzylol wants to merge 2 commits into
stack/528-11-deployment-inputsfrom
stack/528-12-latency

Conversation

@zzylol

@zzylol zzylol commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Problem: query latency bounds are never checked against summary lifecycle alternatives

#509 §3 Plan selection:

Selection rejects every candidate that misses an accuracy target or a latency bound, or that needs a capability the deployment lacks, and then picks the cheapest remaining plan.

#509 §Goal lists latency requirements as part of the query workload, and Example 1 (Stage 3) applies them: "Its cost model estimates Q2's latency against the 100 ms bound; for example, an exact top 10 rebuilt from one million series at every refresh may miss it." #509 §Stages and their decisions also requires that every pruned candidate carries a reason.

Before this PR, the workload already carried the bound, but nothing read it during planning:

// crates/types/src/workload.rs
pub enum LatencyRequirement { ExplicitMaxMs(f64), Unspecified }
pub struct QueryRequirements {
    pub accuracy: AccuracyRequirement,
    pub response_latency: LatencyRequirement, // only workload validation read this
}

CostModel returns abstract Cost / CostRate values and had no latency hook. So for this workload:

Query Repeats Latency requirement
quantile_over_time(0.99, lat[5m]) yes ≤ 100 ms

a deployment whose ephemeral KLL takes 250 ms per read (rebuilt from raw samples) and whose continuously maintained KLL takes 50 ms had no way to say so. The ephemeral alternative stayed selectable whenever it was cheaper, and the 100 ms bound had no effect on the plan.

Scope. This PR covers the latency-bound part of #509 §3 for summary lifecycle alternatives (part 1 of #526). It leaves out: a latency check on raw recomputation (#556), and recording reasons for valid candidates that lose selection (part 2 of #526).

Proposed method

All changes are in the lifecycle stage, which chooses a summary's materialization lifecycle and feeds selection (crates/asap-aware-mapping/src/summary_maintenance_lifecycle.rs).

  1. Deployment hook. Add CostModel::summary_read_latency_ms(summary, lifecycle) -> Option<f64>. It is the deployment's estimate of response latency for reading summary under one physical lifecycle. The default returns None (no estimate).
  2. Strictest bound. workload_facts already walks the workload entries that consume a summary. It now also keeps the minimum ExplicitMaxMs bound over those entries, in the new field SummaryMaintenanceWorkloadFacts.latency_bound_ms. Unspecified entries add nothing. If no entry has a bound, the field is None and no latency check runs.
  3. Per-alternative check. In enumerate_with_profile, after alternatives_for builds the alternatives for one summary, and only when a bound exists, call the hook for each alternative's lifecycle:
    • Estimate finite, >= 0 and > bound: if the alternative has no rejection yet, set rejection = ExceedsLatencyBound and push the assumption "estimated response latency {latency_ms} ms exceeds {bound_ms} ms bound".
    • Estimate finite, >= 0 and <= bound: no change.
    • None, NaN, infinite or negative: keep the alternative and push the assumption "latency bound {bound_ms} ms unchecked: no estimate".
  4. An alternative with a rejection is not selectable (selectable() requires rejection.is_none() and a known cost), so selection moves to the next legal lifecycle. An existing rejection is never overwritten, so the first reason stays.

The check is per lifecycle alternative, not per logical candidate, because the same summary can be fast when maintained and slow when rebuilt per read.

Key code interfaces

crates/asap-aware-mapping/src/cost_model.rs

pub trait CostModel {
    // … existing methods …

    /// Estimated response latency for reading `summary` under one selected
    /// physical lifecycle. `None` means the deployment has no estimate; it
    /// does not make the alternative invalid.
    fn summary_read_latency_ms(
        &self,
        _summary: &OperatorNode,
        _lifecycle: &SummaryMaintenanceLifecycle,
    ) -> Option<f64> {
        None
    }
}

crates/asap-aware-mapping/src/summary_maintenance_lifecycle.rs

pub enum SummaryMaintenanceLifecycleRejection {
    UnsupportedByRuntime,
    // … existing variants …
    MissingCostEvidence,
    ExceedsLatencyBound, // new
}

// private
struct SummaryMaintenanceWorkloadFacts {
    required_accuracy: Vec<AccuracyTarget>,
    latency_bound_ms: Option<f64>, // new
    // … reads, one_time_invocations, evaluation_rate, …
}

The result is visible on the existing public type:

pub struct SummaryMaintenanceLifecycleAlternative {
    pub summary_maintenance_lifecycle: SummaryMaintenanceLifecycle,
    pub total_cost: Option<Cost>,
    pub rejection: Option<SummaryMaintenanceLifecycleRejection>, // may now be ExceedsLatencyBound
    pub assumptions: Vec<String>,                                // may now hold a latency note
}

Fields

CostModel::summary_read_latency_ms

Parameter / return Type Meaning
summary &OperatorNode The materialized SummaryAgg whose read is estimated
lifecycle &SummaryMaintenanceLifecycle The lifecycle alternative being checked: Ephemeral, Prepared { activate_at, retire_at }, Shared { retention } or ContinuouslyMaintained
return Option<f64> Milliseconds for one response. None = no estimate. Only finite values >= 0 are compared; anything else is treated as no estimate.

Implemented by the deployment's cost model. The default (and DefaultCostModel) returns None.

SummaryMaintenanceLifecycleRejection::ExceedsLatencyBound: the deployment's estimate for this lifecycle is larger than the strictest bound among the summary's consumers. Serialized as "exceeds_latency_bound" (rename_all = "snake_case"). Set only by the latency check, and only on an alternative that had no rejection.

SummaryMaintenanceWorkloadFacts.latency_bound_ms (private): Option<f64>, the minimum LatencyRequirement::ExplicitMaxMs over the workload entries in workload_entry_indices for this summary. None when every consumer is Unspecified. Set by workload_facts.

SummaryMaintenanceLifecycleAlternative fields as used here:

  • rejection: None means legal; Some(ExceedsLatencyBound) is the new reason.
  • assumptions: gets one latency string per checked alternative (either the "exceeds" text with the estimate, or the "unchecked: no estimate" text).
  • total_cost, summary_maintenance_lifecycle: unchanged.

Examples

Slow ephemeral loses to fast maintained. Test latency_bound_rejects_slow_ephemeral_summary_and_keeps_fast_maintained_one in crates/planner/tests/summary_sharing.rs:

  • Input: repeating PromQL quantile_over_time(0.99, lat[5m]), ε = 0.01, response_latency = ExplicitMaxMs(100.0). Test cost model FixedCosts { build: 1.0, raw_per_read: 1_000.0, latency_estimates: true }, whose hook returns 250 ms for Ephemeral and 50 ms for every other lifecycle.
  • What happens: the bound for the KLL summary is 100 ms. Ephemeral (250 > 100) is rejected. The other lifecycles (50 ≤ 100) are left as they are.
  • Output: the deployment's alternatives contain Ephemeral with rejection == Some(ExceedsLatencyBound) and ContinuouslyMaintained with rejection == None.

No estimate keeps the candidate. Test missing_latency_estimate_keeps_candidate_and_records_unchecked_reason, same file:

  • Input: same query and 100 ms bound, cost model CHEAP_SUMMARY (latency_estimates: false, so the hook returns None).
  • Output: planning succeeds, and some alternative has rejection == None and an assumption containing "latency bound 100 ms unchecked".
Bound Estimate for the lifecycle Alternative after the check
none (Unspecified) any unchanged; hook not called
100 ms 50 ms unchanged
100 ms 250 ms, no earlier rejection ExceedsLatencyBound + "exceeds" assumption
100 ms 250 ms, already rejected (e.g. UnsupportedByRuntime) keeps the earlier rejection
100 ms None, NaN, ∞ or negative kept; "unchecked: no estimate" assumption
consumers with 200 ms and 100 ms bounds share one summary (rule, no dedicated test) — bound used is 100 ms

The existing tests that build FixedCosts set latency_estimates: false, so their results do not change.

Out of scope

Stack and validation

Legacy physical stack: … ← #553 ← #554 ← #555 ← #556 … · Base: #553 (stack/528-11-deployment-inputs) · Next: #555 · Tracker: #528

Implements the latency part of #526. Selection-loss explanations remain a separate follow-up.

Validation:

  • CARGO_TARGET_DIR=/mydata/cargo-target-412 cargo test --locked -p asap-planner
  • cargo fmt --all -- --check

🤖 Generated with Claude Code

@zzylol

zzylol commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Parked as draft: PR priorities changed (see #528). Order is now (A) finish #511 operator sharing, (B) the #572 crate/module reorganization, (C) #509 end-to-end stages. This PR sits on the old stack/528-legacy-physical-base chain, and Phase B moves the files it touches. Its content will be re-scoped onto the new layout in Phase C.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant