Skip to content

fix(planner): enforce response bounds on raw fallback - #556

Draft
zzylol wants to merge 1 commit into
stack/528-13-summary-mergefrom
stack/528-14-raw-latency
Draft

zzylol wants to merge 1 commit into
stack/528-13-summary-mergefrom
stack/528-14-raw-latency

Conversation

@zzylol

@zzylol zzylol commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Problem: raw recomputation bypasses the response-latency bound

#509 §3 Plan selection:

Selection rejects every candidate that misses an accuracy target or a latency bound, or that needs a capability the deployment lacks, and then picks the cheapest remaining plan.

The rule applies to every candidate, including the exact (no-summary) one. #509 Example 1, Stage 3 uses exactly this case: "Its cost model estimates Q2's latency against the 100 ms bound; for example, an exact top 10 rebuilt from one million series at every refresh may miss it." In #509 Example 4 Pattern B, B2 (rescan raw data on each read) is one of the candidates selection compares.

#554 added the latency check for summary lifecycle alternatives only. Raw recomputation was still compared on amortized cost alone. Two failures follow, both with this workload and a deployment cost model:

Query Repeats Latency requirement
quantile_over_time(0.99, lat[5m]) yes ≤ 100 ms
Alternative Cost Response latency
raw recompute 0.01 per read 250 ms
KLL, Ephemeral build 100 250 ms
KLL, other lifecycles build 100 50 ms
  1. Cheap raw wins. Raw is far cheaper over the horizon, so selection picks it, even though 250 ms misses the 100 ms bound and a 50 ms maintained summary exists.
  2. Silent fallback. If the runtime supports only ephemeral state, feat(planner): enforce summary response latency bounds #554 rejects the only summary alternative (ExceedsLatencyBound). Final assembly then falls back to raw recomputation, which is just as slow, and planning succeeds with a plan that misses the bound.

Scope. This PR completes the latency-bound part of #509 §3 for the raw (exact) alternative in lifecycle selection. It leaves out: per-candidate reasons for valid plans that lose selection (#526 part 2), latency of anything other than the original query executed once, and any built-in latency estimate.

Proposed method

All changes are in the lifecycle stage (crates/asap-aware-mapping/src/summary_maintenance_lifecycle.rs), at the two points where raw recomputation is compared with summaries.

  1. Deployment hook. Add CostModel::raw_query_response_latency_ms(target) -> Option<f64>: the latency of executing the original query once. It is separate from raw_query_recompute_total_cost, because a plan that is cheap over the horizon can still miss a per-response deadline. Default None.
  2. Violation check. New private raw_response_latency_violation(target, demand, cost_model) -> Option<(bound_ms, estimate_ms)>:
    • bound = the minimum ExplicitMaxMs over the workload entries in demand.entry_indices; if none, return None (no check);
    • estimate = the hook's value; if None, return None (unchecked, as in feat(planner): enforce summary response latency bounds #554);
    • violation when the estimate is greater than the bound, or is not finite, or is negative.
  3. Global selection (global_selection_with_summary_maintenance_lifecycles). Before, the comparison was atomic: if raw cost was unknown, no summary override was recorded. Now:
    • raw admitted (no violation) and raw cost known: insert raw cost and the summary cost, as before;
    • raw cost unknown and raw admitted: nothing recorded, as before;
    • raw violates the bound: raw cost is not inserted, and the summary cost is still recorded, so the summary candidate can win without a raw comparison.
  4. Final assembly (plan_assembled_dag). If raw violates the bound:
    • and the plan selected raw recompute, or no summary alternative has a cost (summary_total_cost is None): return NoLatencyFeasiblePlan { bound_ms, estimate_ms };
    • otherwise keep the summary plan. The existing "raw is cheaper, switch to raw" step is skipped when raw violates the bound.
  5. With no raw estimate, behavior is unchanged: the bound is unchecked for raw, matching feat(planner): enforce summary response latency bounds #554's policy for summaries.

Key code interfaces

crates/asap-aware-mapping/src/cost_model.rs

pub trait CostModel {
    // … existing methods, including raw_query_recompute_cost / raw_query_recompute_total_cost …

    /// Response latency for executing the original ordinary query once.
    /// Separate from amortized workload cost: cheap recomputation can still
    /// miss a response deadline. `None` means the bound cannot be checked.
    fn raw_query_response_latency_ms(&self, _target: &OperatorNode) -> Option<f64> {
        None
    }
}

crates/asap-aware-mapping/src/summary_maintenance_lifecycle.rs

#[derive(Debug, thiserror::Error)]
pub enum SummaryMaintenanceLifecyclePlanError {
    #[error("no latency-feasible plan: raw response estimate {estimate_ms} ms cannot meet {bound_ms} ms, and no costed summary alternative is available")]
    NoLatencyFeasiblePlan { bound_ms: f64, estimate_ms: f64 }, // new
    #[error(transparent)]
    InvalidWorkload(#[from] WorkloadError),
    // … existing variants …
}

// private
fn raw_response_latency_violation(
    target: &OperatorNode,
    demand: WorkloadDemand<'_>,
    cost_model: &dyn CostModel,
) -> Option<(f64, f64)>; // (bound_ms, estimate_ms)

plan_assembled_dag returns SummaryMaintenanceLifecycleAssemblyError; the new variant reaches it through .into(), and MajorPass reports it as OptimizeError::LifecycleAssembly.

Fields

CostModel::raw_query_response_latency_ms

Parameter / return Type Meaning
target &OperatorNode The original (ordinary, no-summary) query sub-DAG that raw recomputation would execute
return Option<f64> Milliseconds for one execution. None = no estimate, bound unchecked. A non-finite or negative value counts as a violation.

Implemented by the deployment's cost model. The default (and DefaultCostModel) returns None.

SummaryMaintenanceLifecyclePlanError::NoLatencyFeasiblePlan

Field Type Meaning
bound_ms f64 Strictest ExplicitMaxMs among the target's consuming workload entries
estimate_ms f64 The raw latency estimate that violated it

Returned only by final assembly, when raw violates the bound and no costed summary alternative is left (or raw was selected).

raw_response_latency_violation (private)

Parameter / return Meaning
target Passed to the hook
demand: WorkloadDemand workload + data_workload + entry_indices; only workload and entry_indices are read, to find the bound
cost_model Source of the estimate
return Some((bound_ms, estimate_ms)) Raw is excluded for this target
return None No bound, no estimate, or estimate within the bound

Examples

Both tests are in crates/planner/tests/summary_sharing.rs and use the workload above. Test cost model: FixedCosts { build: 100.0, raw_per_read: 0.01, latency_estimates: true }. Its raw_query_response_latency_ms returns 250 ms; its summary_read_latency_ms (from #554) returns 250 ms for Ephemeral and 50 ms otherwise.

1. Cheap slow raw cannot win. slow_cheap_raw_recompute_cannot_bypass_the_response_bound:

  • Input: default capabilities (ALL).
  • What happens: raw 250 ms > 100 ms, so global selection does not insert the raw cost and final assembly does not switch to raw. Ephemeral is rejected by feat(planner): enforce summary response latency bounds #554; a 50 ms lifecycle stays legal.
  • Output: selected_raw_recompute == false and summary_total_cost.is_some().

2. Only raw is left. slow_raw_only_query_reports_no_latency_feasible_plan:

  • Input: PlanningModels::builtin().with_cost(&costs).with_capabilities(...) with only supports_ephemeral = true.
  • What happens: Ephemeral (250 ms) is rejected, so no summary alternative has a cost. Raw (250 ms) violates the bound.
  • Output: e2e_plan returns an error whose text contains "no latency-feasible plan".
Raw estimate vs 100 ms bound Costed summary alternative left? Result
no bound on any consumer any unchanged; no raw check
None any unchanged; raw unchecked
50 ms any unchanged; normal cost comparison
250 ms yes summary selected, even if raw is cheaper (test 1)
250 ms no NoLatencyFeasiblePlan (test 2)
NaN, ∞ or negative — treated as a violation (code rule; no test)

Note the asymmetry with #554: an invalid summary estimate is treated as "no estimate", while an invalid raw estimate is treated as a violation.

Out of scope

Stack and validation

Legacy physical stack: … ← #554 ← #555 ← #556 ← #557 … · Base: #555 (stack/528-13-summary-merge) · Next: #557 · Tracker: #528

Part of #526 (latency bounds; follows #554). Documented in docs/develop_docs/operator-design-acceptance.md ("Planner-layering follow-up: raw response latency").

Validation: both new public-planner E2E regressions fail before the fix and pass afterward; planner and mapping suites (524 tests passed); formatting; Clippy for planner and mapping, all targets, warnings denied.

🤖 Generated with Claude Code

@zzylol

zzylol commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Parked as draft: PR priorities changed (see #528). Order is now (A) finish #511 operator sharing, (B) the #572 crate/module reorganization, (C) #509 end-to-end stages. This PR sits on the old stack/528-legacy-physical-base chain, and Phase B moves the files it touches. Its content will be re-scoped onto the new layout in Phase C.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant