Skip to content

feat(planner): whole-expression top-k heap sketch over raw samples - #589

Draft
zzylol wants to merge 1 commit into
stack/509-e1-pass2-identicalfrom
stack/509-e2-keyed-additive-topk
Draft

zzylol wants to merge 1 commit into
stack/509-e1-pass2-identicalfrom
stack/509-e2-keyed-additive-topk

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Rebased on main d4869a7 (DF 54). At this head cargo fmt --all --check, cargo clippy --workspace --all-targets -- -D warnings and cargo test --workspace pass (1517 passed, 0 failed, 10 ignored).

Why

For Q2 topk by (job) (10, sum_over_time(http_requests_total[1m])), Pass 1 offered only a heap sketch over an exact per-series sum_over_time: 1M per-series sums feeding the sketch. The whole expression can instead be one CountSketch+heap (or CMS+heap) per job that reads the raw samples. The item is the series identity and the weight is the sample value. This is the legacy keyed-additive realization. It is also the tree DP's "outer choice drops the inner target" case, which had never been exercised. Stacked on #588. Refs #509, #580, #572.

What

  • Pass 1 (logical_candidates.rs). A top-k over a per-item sum/count whose inner aggregate has no other consumer gets whole-expression Count-Min and CountSketch + heap alternatives. LocalLogicalTarget::absorbs records the inner target each one absorbs. Applicability, weight and weight domain come from the existing realize_keyed_additive_summary_input. The item is the series-identity column when rows carry it, as in the existing top-k path.
  • Enumeration. An absorbed target has no choice: enumerate_choices skips its non-pass-through choices, combination_count counts the valid choices, and choice_index ranks in that order, so ids stay contiguous.
  • DP (plan-selection). The DP did not handle this unchanged. It added best(u) for every alternative of t, so an absorbing alternative would have been credited with the inner target's own saving. Now best(t,c) = local(t,c) + Σ best(u) runs only over read_targets(t,c). The pair guard skips pairs an alternative does not read, and the absorbed target is set to its pass-through.
  • Pricing fix the guard found. Top-k readout rows were min(rows the summary read, k × groups). A sketch over raw samples reads several rows per item, so the readout was sized differently from the exact path. On count(topk(…)) with 3 series this coupled with the outer count and forced exhaustive fallback. Readout rows are now bounded by the series count too.
  • Accuracy/validity unchanged. Count-Min still needs non-negative weights. Raw http_requests_total samples have no proof (UnitCount / ResetAwareCounterDerivative do not apply), so whole-expression CMS+heap is kept in Stage 1 and rejected by Stage 3 and by the runtime for the same reason. CountSketch admits signed weights. Its guarantee is the sketch's own analytical L2 bound on the whole expression.
  • Devtool label: "Q2 whole-expression CountSketch+heap".

How it was checked

  • Pass 1 unit test: 8 valid choices for this Q2, choice_index equals the enumeration position, and the absorbed sum_over_time is not in the composed DAG.
  • Plan-selection unit test: with real costs, and with a bonus that makes the whole-expression alternative cheapest, the DP equals brute force over every valid choice and sets the inner target to pass-through.
  • DP = exhaustive on Example 1 (64 combinations, P58) and on count(topk(…)) (40) at 3 and 1M series.
  • The runtime compiles every valid candidate, including the whole-expression CountSketch ones (stage2_runtime_compiles_exactly_the_candidates_stage3_finds_valid, stage2_count_sketch_heap_topk_compiles_in_the_physical_planner).
  • New q2_roots_keep_one_schema_across_realizations: every Q2 option has the selected-rows schema (job, series identity, value) in Stage 1.

Before this PR (Example 1, after #588)

1 → 48 → 48 → 1, with 16 invalid (Count-Min + heap). Selected P44 "Q1 exact (Sum acc, Rate acc) · Q2 exact (Sum acc) · shared input", 52.201. The only Q2 CountSketch plans sketch the 1M exact per-series sums; the best is 164.201.

After this PR (Example 1)

1 → 64 → 64 → 1 (32 per sharing variant; 24 invalid: all Count-Min + heap, including whole-expression). Selected P58, the same plan as before (renumbered), at 52.201. Stage 3 ranks Q2's options (shared input, Q1 with both accumulators):

Q2 realization Q2 nodes Total
exact: Sum acc (4.0 + 1.0) + sort 1M rows / 100 partitions (14.0) + limit (0.001) 19.001 52.201 (P58, selected)
CountSketch+heap over the exact per-series sums: Sum acc 5.0 + sketch 1M × 126 = 126.0 + estimate 0.001 131.001 164.201 (P62)
whole-expression CountSketch+heap: sketch 4M raw samples × 126 = 504.0 + estimate 0.001; no sum_over_time 504.001 537.201 (P64)

Shared nodes in all three: scan 23.2, range 4.0, Q1 Rate acc 5.0, Q1 Sum acc 1.0. Constants are not tuned. The sketch reads 4× more rows (4M samples vs 1M per-series sums) at depth 125 + 1 heap update per row, so exact wins.

CountSketch sizing (accuracy/estimators/count_sketch.rs): width ceil(3/ε²) = 30,000. Depth is the smallest odd integer ≥ 18·ln(1/δ) (median-of-rows Hoeffding bound exp(−depth/18) ≤ δ): 18·ln(1000) = 124.3 → 125. Stage 3 prices a heap sketch at depth + 1 = 126 ops per row.

Gate

cargo fmt --all --check ✓; cargo clippy --workspace --all-targets --all-features -- -D warnings ✓; cargo test --workspace: 1,517 passed / 10 ignored (#587: 1,508 / 12; #588: 1,514 / 10). The 3 new tests are the Pass 1 whole-expression test, the drop-inner DP test and the Q2 schema test. Viewer: python3 -m unittest discover -s tools/dag-viewer -p test_render.py, 29 OK (6 skipped).

Remaining gaps

  • Stage 2 still exports exact top-k as sort → limit over per-series rows, so the physical Q2 root schema of exact plans carries ts and differs from the sketch readout's (job, identity, value). This predates this PR; the Stage 1 schemas agree.
  • Hydra is still missing (stage1_q2_summary_families_are_heap_sketches_and_hydra stays ignored).

🤖 Generated with Claude Code

Pass 1 now offers, for a top-k over a per-item sum or count whose inner
aggregate has no other consumer, Count-Min and CountSketch + heap
alternatives that read the inner aggregate's input and absorb it (the
legacy keyed-additive rule decides applicability and the update). For
#509 Example 1's Q2 that is one heap per job over the raw samples, keyed
by series identity and weighted by the sample value.

An absorbed target has no choice of its own: enumeration skips its
choices, choice_index ranks in that order, and the tree DP sums only the
targets an alternative still reads (read_targets) and sets the absorbed
target to its pass-through. Top-k readout rows are bounded by the series
count, so a sketch reading several samples per item is sized like the
other realizations (the coupling guard caught the difference).

Example 1: 1 -> 64 -> 64 -> 1; P58 (all exact, shared input) at 52.201;
the best whole-expression CountSketch + heap plan, P64, costs 537.201.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/509-e1-pass2-identical branch from 41c0c78 to a52409e Compare October 5, 2026 06:21
@zzylol
zzylol force-pushed the stack/509-e2-keyed-additive-topk branch from 6575cfe to 4bcdab4 Compare October 5, 2026 06:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant