From fe519ae12db8969087f6a5be9b36b0873c6e3b97 Mon Sep 17 00:00:00 2001 From: Zeying Zhu Date: Tue, 5 May 2026 15:29:35 -0400 Subject: [PATCH] eval(deploy): mirror inference patterns + align queries-e2e to engine matcher MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Mirrors `ASAPQuery-backend` PR #79's 33-entry warm-tier PromQL pattern set into the 5 deploy `backend-inference{,-cms,-cs,-hll,-kll}.yaml` overlays, and swaps `deploy/scripts/queries-e2e.json` off the engine- rejected `histogram_quantile(...)` shape onto the live-verified `quantile_over_time(φ, *_quantile[1m])` shape. Pre/post coverage per overlay (entry count): | YAML | Before | After | |-------------------------------|--------|-------| | backend-inference.yaml | 8 | 33 | | backend-inference-cms.yaml | 2 | 14 | | backend-inference-cs.yaml | 2 | 16 | | backend-inference-hll.yaml | 3 | 8 | | backend-inference-kll.yaml | 3 | 16 | Per-sketch overlays are filtered subsets (CMS/CountSketch keep Sum/Count/Topk/{sum,count}_over_time/rate/increase; HLL keeps Cardinality/count_over_time; KLL keeps quantile/quantile_over_time). Metric names match each overlay's agent-side `metric_suffix`: `_quantile` for the default DDSketch path (gateway-aggregate-from-raw), `_hll` / `_kll` for the HLL / KLL direct overlays, raw `http_requests_total` for CMS / CountSketch (which preserve metric_name). queries-e2e.json before → after: - histogram_quantile(0.99, sum by (le) (..._latency_ms)) → quantile_over_time(0.99, ..._latency_ms_quantile[1m]) - histogram_quantile(0.50, sum by (le) (..._latency_ms)) → quantile_over_time(0.5, ..._latency_ms_quantile[1m]) - topk(10, http_requests_total) → topk(10, http_requests_total) (kept) - count(count by (zone) (http_requests_total)) → count(http_requests_total) (single-level, matches PR #79 spatial Count pattern) - sum(http_requests_total) → sum_over_time(http_requests_total[1m]) (matches PROGRESS.md 2026-04-30 live-verified CMS/CountSketch path) The engine pattern-matcher (`asap-query-engine/src/engines/simple_engine.rs::controller_patterns`) does NOT include a `histogram_quantile` block — that requires a separate engine PR. Documented in `deploy/README.md` "Inference dispatch" section as the open gap. Verification: * All 5 YAMLs load cleanly through the runtime `read_inference_config` parser (entry counts above). * Every PromQL string in every YAML parses through `promql_parser`. * Every queries-e2e.json query exact-matches an entry in the main `backend-inference.yaml` (so `find_query_config` will hit warm tier, not fall through to cold). Closes the YAML side of E0 exit criterion (1). Co-Authored-By: Claude Opus 4.7 (1M context) --- PROGRESS.md | 15 ++ deploy/README.md | 58 ++++++++ deploy/configs/backend-inference-cms.yaml | 77 +++++++++- deploy/configs/backend-inference-cs.yaml | 92 +++++++++++- deploy/configs/backend-inference-hll.yaml | 43 ++++++ deploy/configs/backend-inference-kll.yaml | 75 ++++++++++ deploy/configs/backend-inference.yaml | 173 +++++++++++++++++++++- deploy/scripts/queries-e2e.json | 8 +- 8 files changed, 523 insertions(+), 18 deletions(-) diff --git a/PROGRESS.md b/PROGRESS.md index 368cbd78..ddd8c9ab 100644 --- a/PROGRESS.md +++ b/PROGRESS.md @@ -2,6 +2,21 @@ _Last updated: 2026-05-05._ +## Deploy-side inference mirror + queries-e2e alignment (2026-05-05) + +Mirrored `ASAPQuery-backend` PR #79's 33-entry warm-tier pattern set +into the 5 deploy `backend-inference{,-cms,-cs,-hll,-kll}.yaml` +overlays (was 1–8 entries each; now 33 / 14 / 16 / 8 / 16) and +swapped `deploy/scripts/queries-e2e.json` off the engine-rejected +`histogram_quantile(...)` shape onto `quantile_over_time(φ, +*_quantile[1m])` (live-verified per "Single-pipeline multi-sketch + +delta + queryable warm tier (2026-05-01)" above). Closes the YAML +side of the E0 exit criterion (1) — replay queries now exact-match +warm-tier entries instead of falling through to the cold tier. The +engine pattern-matcher itself does not yet cover `histogram_quantile`; +that's a separate ASAPQuery-backend PR (see +`deploy/README.md` "Inference dispatch"). + ## Cross-language byte-format parity, 5/5 sketches (2026-05-05) Closes [#243](https://github.com/ProjectASAP/ASAPCollector/issues/243). diff --git a/deploy/README.md b/deploy/README.md index 08973bb2..2fc49f0a 100644 --- a/deploy/README.md +++ b/deploy/README.md @@ -85,6 +85,59 @@ docker compose -f deploy/docker-compose/base.yml \ -f deploy/docker-compose/baseline-b3-delta.yml up ``` +### Inference dispatch (warm-tier query coverage) + +The backend's "warm tier" (precomputed-sketch path) only answers +queries whose canonical PromQL string exact-matches an entry in the +mounted inference YAML. Five overlays live in `deploy/configs/`: + +| YAML | Mounted by | Covers (PromQL families × ranges) | Entries | +|---|---|---|---| +| `backend-inference.yaml` | `e2e-overlay.yml`, `queryengine-overlay.yml` (default) | All 33 patterns from `ASAPQuery-backend` PR #79: spatial multi-quantile, `quantile_over_time(φ ∈ {0.5, 0.9, 0.95, 0.99}, …[1m\|2m\|5m])`, `sum_over_time` / `count_over_time` × wider ranges, `rate` / `increase`, spatial `count` / `sum` / `avg`, `topk(5\|10\|50, …)`. | 33 | +| `backend-inference-cms.yaml` | `e2e-overlay-cms.yml` | CountMinSketch families: `{sum, count, avg}`, `{sum_over_time, count_over_time, rate, increase}` × `[1m, 2m, 5m]`. | 14 | +| `backend-inference-cs.yaml` | `e2e-overlay-cs.yml` | CountSketch families: same as CMS plus `topk(5\|10\|50, …)`. | 16 | +| `backend-inference-hll.yaml` | `e2e-overlay-hll.yml` | HLL cardinality families: spatial `count(metric_hll)` and `count_over_time(metric_hll[1m\|2m\|5m])` for both the counter (`http_requests_total_hll`) and gauge (`http_requests_total_latency_ms_hll`) flavours. | 8 | +| `backend-inference-kll.yaml` | `e2e-overlay-kll.yml` | KLL rank-quantile families: spatial `quantile by (zone) (φ, …)` and `quantile_over_time(φ, metric_kll[1m\|2m\|5m])` × `φ ∈ {0.5, 0.9, 0.95, 0.99}`. | 16 | + +This is the canonical paper-experiment pattern set (PR #79 +`tests/inference_yaml_pattern_coverage.rs` is the runtime contract). +Adding a new query family here without a matching entry in PR #79's +canonical YAML risks shipping warm-tier "promises" the engine can't +keep — the YAML is checked exact-string at request time by +`find_query_config`. + +#### Metric-name conventions per overlay + +| Overlay | Backend-side metric name | Why | +|---|---|---| +| `backend-inference.yaml` | `http_requests_total_latency_ms_quantile` (KLL/DDSketch quantile patterns), `http_requests_total` (CMS/CountSketch/HLL Sum/Count/Topk patterns) | Default deploy uses `gateway-aggregate-from-raw.yaml` → DDSketch → `metric_suffix: "_quantile"`. Sum/Count/Topk patterns target the raw counter forwarded unsuffixed by CMS/CountSketch processors (see `sketchcol-agent-{cms,cs}-direct.yaml`). | +| `backend-inference-cms.yaml`, `-cs.yaml` | `http_requests_total` | CMS / CountSketch direct agents preserve the raw metric name (no `metric_suffix`). | +| `backend-inference-hll.yaml` | `http_requests_total_hll`, `http_requests_total_latency_ms_hll` | `sketchcol-agent-hll-direct.yaml` adds `metric_suffix: "_hll"`. | +| `backend-inference-kll.yaml` | `http_requests_total_latency_ms_kll` | `sketchcol-agent-kll-direct.yaml` adds `metric_suffix: "_kll"`. | + +#### Open gap: `histogram_quantile(φ, …)` is NOT covered + +The engine pattern matcher +(`asap-query-engine/src/engines/simple_engine.rs::controller_patterns`) +includes `quantile_over_time` and the spatial `quantile by (…)` ops +but does **not** include a `histogram_quantile` pattern block. +`histogram_quantile(0.99, sum by (le) (http_requests_total_latency_ms))` +fails the engine matcher before reaching `find_query_config`, even +when an entry of that exact string is present in the YAML. + +Two paths forward (in priority order): + +1. **Today (this PR's choice for `queries-e2e.json`):** use the + pre-aggregated `quantile_over_time(φ, *_quantile[…])` shape. The + gateway / agent path produces `http_requests_total_latency_ms_quantile` + (DDSketch / KLL backed) which the engine matches and answers. + PROGRESS.md "Single-pipeline multi-sketch + delta + queryable + warm tier (2026-05-01)" verified this path live (`q=0.5 → + 19.49`). +2. **Tomorrow (separate engine PR):** extend `controller_patterns` + with a `histogram_quantile` block. Out of scope for E0; tracked + alongside PR #79 follow-ups. + ### E0: single-cell smoke (P5–P9 end-to-end against a live stack) Smallest cell that exercises the whole P1–P9 path. Use this @@ -126,6 +179,11 @@ python3 deploy/scripts/plan_transition.py \ --transition-out /tmp/cell-smoke-e0/transition.jsonl \ --sample-out /tmp/cell-smoke-e0/sample.jsonl \ --soak-secs 60 --pre-transition-secs 20 & +# Note: the `histogram_quantile(...)` shape above is INTENTIONALLY +# unmatched by the engine — it's the capability-miss probe used by +# `plan_transition.py` to drive a fresh plan publish. Replay-side +# queries (`queries-e2e.json`) deliberately use shapes that DO match +# (see "Inference dispatch" above). wait # 4. Snapshot ground truth from the cold-store volume (the volume diff --git a/deploy/configs/backend-inference-cms.yaml b/deploy/configs/backend-inference-cms.yaml index b682535b..11509aa7 100644 --- a/deploy/configs/backend-inference-cms.yaml +++ b/deploy/configs/backend-inference-cms.yaml @@ -1,3 +1,27 @@ +# Warm-tier dispatch table (CountMinSketch single-sketch overlay). +# +# Subset of `backend-inference.yaml` (which mirrors PR #79's +# `inference_config.yaml`) restricted to the families that route to +# the CountMinSketch accumulator: +# +# * Spatial `sum` / `count` — Statistic::{Sum,Count} no-key +# fallback returns the min-row sum (canonical CMS total-event +# estimator). PROGRESS.md "All-five-sketch runtime e2e +# verification (2026-04-30)" verified +# `sum_over_time(http_requests_total[1m])` → 145735 with ε≈0.0027. +# * `sum_over_time` / `count_over_time` over [1m]/[2m]/[5m] — +# Statistic::Sum / Statistic::Count via +# `additive_frequency` accuracy envelope. +# * `rate` / `increase` — same family (delta arithmetic over Sum). +# +# `topk` / `histogram_quantile` / `quantile_over_time` excluded +# (CMS does not back those statistics — see the +# `compatible_agg_types` table in `capability_matching.rs`). +# +# Metric: `http_requests_total` (no suffix). The CMS direct agent +# overlay (`sketchcol-agent-cms-direct.yaml`) uses +# `metric_name: "http_requests_total"` and preserves the raw name +# end-to-end. cleanup_policy: name: "circular_buffer" metrics: @@ -7,6 +31,20 @@ metrics: - node - pod queries: +# ── Spatial Sum / Count / Avg ───────────────────────────────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: sum(http_requests_total) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count(http_requests_total) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: avg(http_requests_total) +# ── sum_over_time / count_over_time × wider ranges ──────────────────── - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 @@ -14,4 +52,41 @@ queries: - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 - query: sum(http_requests_total) + query: sum_over_time(http_requests_total[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: sum_over_time(http_requests_total[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[5m]) +# ── rate / increase (delta arithmetic over Sum family) ──────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: increase(http_requests_total[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: increase(http_requests_total[5m]) diff --git a/deploy/configs/backend-inference-cs.yaml b/deploy/configs/backend-inference-cs.yaml index 9d1a7376..c8d9f663 100644 --- a/deploy/configs/backend-inference-cs.yaml +++ b/deploy/configs/backend-inference-cs.yaml @@ -1,3 +1,26 @@ +# Warm-tier dispatch table (CountSketch single-sketch overlay). +# +# Subset of `backend-inference.yaml` (which mirrors PR #79's +# `inference_config.yaml`) restricted to the families that route to +# the CountSketch accumulator: +# +# * `topk(N, …)` — Statistic::Topk; CountSketch's `query_statistic` +# returns the row-mean total when no key is supplied (per PR #79 +# wiring; per-key enumeration needs a paired SetAggregator on the +# agent — tracked upstream). +# * Spatial `sum` / `count` — Statistic::{Sum,Count} no-key +# fallback (PROGRESS.md "All-five-sketch runtime e2e verification +# (2026-04-30)" — `sum_over_time(http_requests_total[1m])` → +# 24266 with ε=0.03 `additive_frequency` envelope). +# * `sum_over_time` / `count_over_time` × [1m]/[2m]/[5m]. +# * `rate` / `increase` — same Sum family. +# +# `quantile_over_time` / `histogram_quantile` excluded (CountSketch +# does not back rank statistics). +# +# Metric: `http_requests_total` (no suffix). The CountSketch direct +# agent overlay (`sketchcol-agent-cs-direct.yaml`) preserves the raw +# metric name through the processor. cleanup_policy: name: "circular_buffer" metrics: @@ -6,17 +29,72 @@ metrics: - rack - node - pod - http_requests_total_latency_ms: - - zone - - rack - - node - - pod queries: +# ── Spatial Sum / Count ─────────────────────────────────────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: sum(http_requests_total) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count(http_requests_total) +# ── sum_over_time / count_over_time × wider ranges ──────────────────── - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 query: sum_over_time(http_requests_total[1m]) - aggregations: - - aggregation_id: 2 + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: sum_over_time(http_requests_total[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: sum_over_time(http_requests_total[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[5m]) +# ── rate / increase (delta arithmetic over Sum family) ──────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: increase(http_requests_total[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: increase(http_requests_total[5m]) +# ── Top-K (CountSketch heavy-hitter readout) ────────────────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: topk(5, http_requests_total) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: topk(10, http_requests_total) +- aggregations: + - aggregation_id: 1 num_aggregates_to_retain: 6 - query: sum_over_time(http_requests_total_latency_ms[1m]) + query: topk(50, http_requests_total) diff --git a/deploy/configs/backend-inference-hll.yaml b/deploy/configs/backend-inference-hll.yaml index 41308f12..2d5b9470 100644 --- a/deploy/configs/backend-inference-hll.yaml +++ b/deploy/configs/backend-inference-hll.yaml @@ -1,3 +1,24 @@ +# Warm-tier dispatch table (HyperLogLog single-sketch overlay). +# +# Subset of `backend-inference.yaml` (which mirrors PR #79's +# `inference_config.yaml`) restricted to the families that route to +# the HLL accumulator: +# +# * `count(metric)` — Statistic::Count; HLL's `query_statistic` +# accepts both `Cardinality` and `Count` (alias) and returns the +# unique-cardinality estimate. PROGRESS.md "All-five-sketch +# runtime e2e verification (2026-04-30)" verified +# `count(http_requests_total)` → 149.68 distinct, ε=0.008 +# `relative_cardinality` envelope. +# * `count_over_time(metric[…])` — temporal Count. +# +# `sum` / `quantile_over_time` / `topk` / `rate` excluded (HLL is a +# cardinality sketch — no value semantics). +# +# Metric: `http_requests_total_hll` / `http_requests_total_latency_ms_hll` +# — the HLL direct agent overlay (`sketchcol-agent-hll-direct.yaml`) +# uses `metric_suffix: "_hll"`, so the backend-side metric name +# is the suffixed form. cleanup_policy: name: "circular_buffer" metrics: @@ -12,15 +33,37 @@ metrics: - node - pod queries: +# ── Spatial Cardinality (count over an HLL-backed metric) ───────────── - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 query: count(http_requests_total_hll) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count(http_requests_total_latency_ms_hll) +# ── count_over_time × wider ranges ──────────────────────────────────── - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 query: count_over_time(http_requests_total_hll[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total_hll[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total_hll[5m]) - aggregations: - aggregation_id: 2 num_aggregates_to_retain: 6 query: count_over_time(http_requests_total_latency_ms_hll[1m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total_latency_ms_hll[2m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total_latency_ms_hll[5m]) diff --git a/deploy/configs/backend-inference-kll.yaml b/deploy/configs/backend-inference-kll.yaml index fc510f21..46e3903c 100644 --- a/deploy/configs/backend-inference-kll.yaml +++ b/deploy/configs/backend-inference-kll.yaml @@ -1,3 +1,24 @@ +# Warm-tier dispatch table (KLL single-sketch overlay). +# +# Subset of `backend-inference.yaml` (which mirrors PR #79's +# `inference_config.yaml`) restricted to the families that route to +# the KLL accumulator: +# +# * `quantile_over_time(φ, metric[range])` — Statistic::Quantile; +# KLL is the rank-quantile sketch (datasketches-cpp KLL). PR #79 +# `tests/inference_yaml_pattern_coverage.rs::quantile_over_time_multi_phi_routes_through_warm_tier` +# verifies this routing end-to-end. +# * Spatial `quantile by (label) (φ, metric)` — Statistic::Quantile +# spatial-of-temporal. +# +# `sum` / `count` / `rate` / `increase` / `topk` excluded (KLL backs +# rank statistics, not additive frequency). Note: KLL has NO delta +# encoding, so the `rate` / `increase` family in the canonical PR #79 +# pattern set is paired with an Increase-typed accumulator, not KLL. +# +# Metric: `http_requests_total_latency_ms_kll`. The KLL direct agent +# overlay (`sketchcol-agent-kll-direct.yaml`) uses +# `metric_suffix: "_kll"`. cleanup_policy: name: "circular_buffer" metrics: @@ -7,6 +28,24 @@ metrics: - node - pod queries: +# ── Spatial quantile (multi-quantile) ───────────────────────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.5, http_requests_total_latency_ms_kll) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.9, http_requests_total_latency_ms_kll) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.95, http_requests_total_latency_ms_kll) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.99, http_requests_total_latency_ms_kll) +# ── quantile_over_time: multi-quantile × wider ranges ───────────────── - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 @@ -15,7 +54,43 @@ queries: - aggregation_id: 1 num_aggregates_to_retain: 6 query: quantile_over_time(0.9, http_requests_total_latency_ms_kll[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.95, http_requests_total_latency_ms_kll[1m]) - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 query: quantile_over_time(0.99, http_requests_total_latency_ms_kll[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.5, http_requests_total_latency_ms_kll[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.9, http_requests_total_latency_ms_kll[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.95, http_requests_total_latency_ms_kll[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.99, http_requests_total_latency_ms_kll[2m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.5, http_requests_total_latency_ms_kll[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.9, http_requests_total_latency_ms_kll[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.95, http_requests_total_latency_ms_kll[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.99, http_requests_total_latency_ms_kll[5m]) diff --git a/deploy/configs/backend-inference.yaml b/deploy/configs/backend-inference.yaml index 33e32ceb..9827a901 100644 --- a/deploy/configs/backend-inference.yaml +++ b/deploy/configs/backend-inference.yaml @@ -1,3 +1,36 @@ +# Warm-tier dispatch table (all-sketch default). +# +# Mirrors `ASAPQuery-backend/asap-query-engine/examples/promql/inference_config.yaml` +# (PR #79 — eval(inference): expand warm-tier PromQL pattern coverage) +# entry-for-entry, transposed onto the metric names the deploy +# stack actually emits: +# +# * `http_requests_total_latency_ms_quantile` — DDSketch / KLL, +# produced by `gateway-aggregate-from-raw.yaml` (`metric_suffix: +# "_quantile"`) from the fake-exporter's `*_latency_ms` gauge. +# * `http_requests_total_quantile` — DDSketch / KLL, +# produced by the same gateway path from the `http_requests_total` +# counter. +# * `http_requests_total` — raw counter +# emitted by fake-exporter and forwarded unsuffixed by the +# CountMinSketch / CountSketch agent processors (which preserve +# metric_name). Routes to CMS / CountSketch on Sum / Count. +# +# Each (query, aggregation_id) pair binds a PromQL pattern to a +# precompute plan. The exact-string match in `find_query_config` +# requires the request query to canonicalize to one of the listed +# strings — wider-range / wider-shape variants need their own entries +# here, otherwise the request falls through to capability matching +# and (failing that) the cold tier. See PR #79's +# `tests/inference_yaml_pattern_coverage.rs` for the runtime contract. +# +# NOT covered: `histogram_quantile(φ, …)`. The engine's +# `controller_patterns` (asap-query-engine/src/engines/simple_engine.rs) +# does not include a `histogram_quantile` pattern block; queries of +# that shape are rejected before reaching this YAML. Use +# `quantile_over_time(φ, *_quantile[…])` against the pre-aggregated +# `*_quantile` metric instead. Tracked upstream for engine +# pattern-matcher extension. cleanup_policy: name: "circular_buffer" metrics: @@ -11,7 +44,32 @@ metrics: - rack - node - pod + http_requests_total: + - zone + - rack + - node + - pod queries: +# ── Spatial quantile (multi-quantile) ───────────────────────────────── +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.5, http_requests_total_latency_ms_quantile) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.9, http_requests_total_latency_ms_quantile) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.95, http_requests_total_latency_ms_quantile) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile by (zone) (0.99, http_requests_total_latency_ms_quantile) +# ── quantile_over_time: multi-quantile × wider ranges ───────────────── +# Routes to `Statistic::Quantile`; supported by DDSketch / KLL +# accumulators. - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 @@ -20,27 +78,130 @@ queries: - aggregation_id: 1 num_aggregates_to_retain: 6 query: quantile_over_time(0.9, http_requests_total_latency_ms_quantile[1m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.95, http_requests_total_latency_ms_quantile[1m]) - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 query: quantile_over_time(0.99, http_requests_total_latency_ms_quantile[1m]) - aggregations: - - aggregation_id: 2 + - aggregation_id: 1 num_aggregates_to_retain: 6 - query: quantile_over_time(0.99, http_requests_total_quantile[1m]) + query: quantile_over_time(0.5, http_requests_total_latency_ms_quantile[2m]) - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 - query: histogram_quantile(0.99, http_requests_total_latency_ms_quantile) + query: quantile_over_time(0.9, http_requests_total_latency_ms_quantile[2m]) - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 - query: histogram_quantile(0.5, http_requests_total_latency_ms_quantile) + query: quantile_over_time(0.95, http_requests_total_latency_ms_quantile[2m]) - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 - query: sum(http_requests_total_latency_ms_quantile) + query: quantile_over_time(0.99, http_requests_total_latency_ms_quantile[2m]) - aggregations: - aggregation_id: 1 num_aggregates_to_retain: 6 - query: sum_over_time(http_requests_total_latency_ms_quantile[1m]) + query: quantile_over_time(0.5, http_requests_total_latency_ms_quantile[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.9, http_requests_total_latency_ms_quantile[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.95, http_requests_total_latency_ms_quantile[5m]) +- aggregations: + - aggregation_id: 1 + num_aggregates_to_retain: 6 + query: quantile_over_time(0.99, http_requests_total_latency_ms_quantile[5m]) +# ── sum_over_time / count_over_time: wider ranges ───────────────────── +# Routes to `Statistic::Sum` / `Statistic::Count`; supported by +# DDSketch / CountSketch / CountMinSketch accumulators (no-key +# total-volume fallback per PROGRESS.md 2026-04-30 verification). +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: sum_over_time(http_requests_total[1m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: sum_over_time(http_requests_total[2m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: sum_over_time(http_requests_total[5m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[1m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[2m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count_over_time(http_requests_total[5m]) +# ── rate / increase: delta-capable sketches ─────────────────────────── +# Routes to `Statistic::Rate` / `Statistic::Increase`; supported by +# IncreaseAccumulator / MultipleIncreaseAccumulator. KLL has no delta +# and will return an error for these — pair this YAML with an +# Increase-typed streaming aggregation when serving rate/increase +# queries. +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[1m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[2m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: rate(http_requests_total[5m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: increase(http_requests_total[1m]) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: increase(http_requests_total[5m]) +# ── Cardinality / generic spatial aggregations ──────────────────────── +# `count` over an HLL-backed aggregation routes to `Statistic::Count`, +# which the HLL accumulator answers as a unique-cardinality estimate. +# `sum` / `avg` route to `Statistic::Sum` / `(Sum, Count)`. +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: count(http_requests_total) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: sum(http_requests_total) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: avg(http_requests_total) +# ── Top-K (CountSketch heavy-hitter readout) ────────────────────────── +# `topk(N, …)` routes to `Statistic::Topk`; CountSketch's +# `query_statistic` returns the row-mean total when no key is supplied +# (limitation: per-key top-K enumeration needs a paired SetAggregator +# on the agent — tracked upstream). +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: topk(5, http_requests_total) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: topk(10, http_requests_total) +- aggregations: + - aggregation_id: 2 + num_aggregates_to_retain: 6 + query: topk(50, http_requests_total) diff --git a/deploy/scripts/queries-e2e.json b/deploy/scripts/queries-e2e.json index 5da4c6fc..68299d66 100644 --- a/deploy/scripts/queries-e2e.json +++ b/deploy/scripts/queries-e2e.json @@ -1,11 +1,11 @@ [ { "kind": "quantile", - "promql": "histogram_quantile(0.99, sum by (le) (http_requests_total_latency_ms))" + "promql": "quantile_over_time(0.99, http_requests_total_latency_ms_quantile[1m])" }, { "kind": "quantile", - "promql": "histogram_quantile(0.50, sum by (le) (http_requests_total_latency_ms))" + "promql": "quantile_over_time(0.5, http_requests_total_latency_ms_quantile[1m])" }, { "kind": "topk", @@ -13,10 +13,10 @@ }, { "kind": "count_unique", - "promql": "count(count by (zone) (http_requests_total))" + "promql": "count(http_requests_total)" }, { "kind": "sum", - "promql": "sum(http_requests_total)" + "promql": "sum_over_time(http_requests_total[1m])" } ]