Conversation
This was referenced Oct 4, 2026
Draft
Share one UnivMon across distinct, L2 and entropy; document why those readouts stay uncertified
#596
Draft
Pass 1 now offers, for a top-k over a per-item sum or count whose inner aggregate has no other consumer, Count-Min and CountSketch + heap alternatives that read the inner aggregate's input and absorb it (the legacy keyed-additive rule decides applicability and the update). For #509 Example 1's Q2 that is one heap per job over the raw samples, keyed by series identity and weighted by the sample value. An absorbed target has no choice of its own: enumeration skips its choices, choice_index ranks in that order, and the tree DP sums only the targets an alternative still reads (read_targets) and sets the absorbed target to its pass-through. Top-k readout rows are bounded by the series count, so a sketch reading several samples per item is sized like the other realizations (the coupling guard caught the difference). Example 1: 1 -> 64 -> 64 -> 1; P58 (all exact, shared input) at 52.201; the best whole-expression CountSketch + heap plan, P64, costs 537.201. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
force-pushed
the
stack/509-e1-pass2-identical
branch
from
October 5, 2026 06:21
41c0c78 to
a52409e
Compare
zzylol
force-pushed
the
stack/509-e2-keyed-additive-topk
branch
from
October 5, 2026 06:21
6575cfe to
4bcdab4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rebased on main d4869a7 (DF 54). At this head
cargo fmt --all --check,cargo clippy --workspace --all-targets -- -D warningsandcargo test --workspacepass (1517 passed, 0 failed, 10 ignored).Why
For Q2
topk by (job) (10, sum_over_time(http_requests_total[1m])), Pass 1 offered only a heap sketch over an exact per-seriessum_over_time: 1M per-series sums feeding the sketch. The whole expression can instead be one CountSketch+heap (or CMS+heap) perjobthat reads the raw samples. The item is the series identity and the weight is the sample value. This is the legacy keyed-additive realization. It is also the tree DP's "outer choice drops the inner target" case, which had never been exercised. Stacked on #588. Refs #509, #580, #572.What
logical_candidates.rs). A top-k over a per-item sum/count whose inner aggregate has no other consumer gets whole-expression Count-Min and CountSketch + heap alternatives.LocalLogicalTarget::absorbsrecords the inner target each one absorbs. Applicability, weight and weight domain come from the existingrealize_keyed_additive_summary_input. The item is the series-identity column when rows carry it, as in the existing top-k path.enumerate_choicesskips its non-pass-through choices,combination_countcounts the valid choices, andchoice_indexranks in that order, so ids stay contiguous.plan-selection). The DP did not handle this unchanged. It addedbest(u)for every alternative oft, so an absorbing alternative would have been credited with the inner target's own saving. Nowbest(t,c) = local(t,c) + Σ best(u)runs only overread_targets(t,c). The pair guard skips pairs an alternative does not read, and the absorbed target is set to its pass-through.min(rows the summary read, k × groups). A sketch over raw samples reads several rows per item, so the readout was sized differently from the exact path. Oncount(topk(…))with 3 series this coupled with the outer count and forced exhaustive fallback. Readout rows are now bounded by the series count too.http_requests_totalsamples have no proof (UnitCount/ResetAwareCounterDerivativedo not apply), so whole-expression CMS+heap is kept in Stage 1 and rejected by Stage 3 and by the runtime for the same reason. CountSketch admits signed weights. Its guarantee is the sketch's own analytical L2 bound on the whole expression.How it was checked
choice_indexequals the enumeration position, and the absorbedsum_over_timeis not in the composed DAG.count(topk(…))(40) at 3 and 1M series.stage2_runtime_compiles_exactly_the_candidates_stage3_finds_valid,stage2_count_sketch_heap_topk_compiles_in_the_physical_planner).q2_roots_keep_one_schema_across_realizations: every Q2 option has the selected-rows schema(job, series identity, value)in Stage 1.Before this PR (Example 1, after #588)
1 → 48 → 48 → 1, with 16 invalid (Count-Min + heap). Selected P44 "Q1 exact (Sum acc, Rate acc) · Q2 exact (Sum acc) · shared input", 52.201. The only Q2 CountSketch plans sketch the 1M exact per-series sums; the best is 164.201.
After this PR (Example 1)
1 → 64 → 64 → 1 (32 per sharing variant; 24 invalid: all Count-Min + heap, including whole-expression). Selected P58, the same plan as before (renumbered), at 52.201. Stage 3 ranks Q2's options (shared input, Q1 with both accumulators):
sum_over_timeShared nodes in all three: scan 23.2, range 4.0, Q1 Rate acc 5.0, Q1 Sum acc 1.0. Constants are not tuned. The sketch reads 4× more rows (4M samples vs 1M per-series sums) at depth 125 + 1 heap update per row, so exact wins.
CountSketch sizing (
accuracy/estimators/count_sketch.rs): widthceil(3/ε²)= 30,000. Depth is the smallest odd integer ≥18·ln(1/δ)(median-of-rows Hoeffding boundexp(−depth/18) ≤ δ):18·ln(1000)= 124.3 → 125. Stage 3 prices a heap sketch at depth + 1 = 126 ops per row.Gate
cargo fmt --all --check✓;cargo clippy --workspace --all-targets --all-features -- -D warnings✓;cargo test --workspace: 1,517 passed / 10 ignored (#587: 1,508 / 12; #588: 1,514 / 10). The 3 new tests are the Pass 1 whole-expression test, the drop-inner DP test and the Q2 schema test. Viewer:python3 -m unittest discover -s tools/dag-viewer -p test_render.py, 29 OK (6 skipped).Remaining gaps
tsand differs from the sketch readout's(job, identity, value). This predates this PR; the Stage 1 schemas agree.stage1_q2_summary_families_are_heap_sketches_and_hydrastays ignored).🤖 Generated with Claude Code