Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions PROGRESS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,21 @@

_Last updated: 2026-05-05._

## Deploy-side inference mirror + queries-e2e alignment (2026-05-05)

Mirrored `ASAPQuery-backend` PR #79's 33-entry warm-tier pattern set
into the 5 deploy `backend-inference{,-cms,-cs,-hll,-kll}.yaml`
overlays (was 1–8 entries each; now 33 / 14 / 16 / 8 / 16) and
swapped `deploy/scripts/queries-e2e.json` off the engine-rejected
`histogram_quantile(...)` shape onto `quantile_over_time(φ,
*_quantile[1m])` (live-verified per "Single-pipeline multi-sketch +
delta + queryable warm tier (2026-05-01)" above). Closes the YAML
side of the E0 exit criterion (1) — replay queries now exact-match
warm-tier entries instead of falling through to the cold tier. The
engine pattern-matcher itself does not yet cover `histogram_quantile`;
that's a separate ASAPQuery-backend PR (see
`deploy/README.md` "Inference dispatch").

## Cross-language byte-format parity, 5/5 sketches (2026-05-05)

Closes [#243](https://github.com/ProjectASAP/ASAPCollector/issues/243).
Expand Down
58 changes: 58 additions & 0 deletions deploy/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,59 @@ docker compose -f deploy/docker-compose/base.yml \
-f deploy/docker-compose/baseline-b3-delta.yml up
```

### Inference dispatch (warm-tier query coverage)

The backend's "warm tier" (precomputed-sketch path) only answers
queries whose canonical PromQL string exact-matches an entry in the
mounted inference YAML. Five overlays live in `deploy/configs/`:

| YAML | Mounted by | Covers (PromQL families × ranges) | Entries |
|---|---|---|---|
| `backend-inference.yaml` | `e2e-overlay.yml`, `queryengine-overlay.yml` (default) | All 33 patterns from `ASAPQuery-backend` PR #79: spatial multi-quantile, `quantile_over_time(φ ∈ {0.5, 0.9, 0.95, 0.99}, …[1m\|2m\|5m])`, `sum_over_time` / `count_over_time` × wider ranges, `rate` / `increase`, spatial `count` / `sum` / `avg`, `topk(5\|10\|50, …)`. | 33 |
| `backend-inference-cms.yaml` | `e2e-overlay-cms.yml` | CountMinSketch families: `{sum, count, avg}`, `{sum_over_time, count_over_time, rate, increase}` × `[1m, 2m, 5m]`. | 14 |
| `backend-inference-cs.yaml` | `e2e-overlay-cs.yml` | CountSketch families: same as CMS plus `topk(5\|10\|50, …)`. | 16 |
| `backend-inference-hll.yaml` | `e2e-overlay-hll.yml` | HLL cardinality families: spatial `count(metric_hll)` and `count_over_time(metric_hll[1m\|2m\|5m])` for both the counter (`http_requests_total_hll`) and gauge (`http_requests_total_latency_ms_hll`) flavours. | 8 |
| `backend-inference-kll.yaml` | `e2e-overlay-kll.yml` | KLL rank-quantile families: spatial `quantile by (zone) (φ, …)` and `quantile_over_time(φ, metric_kll[1m\|2m\|5m])` × `φ ∈ {0.5, 0.9, 0.95, 0.99}`. | 16 |

This is the canonical paper-experiment pattern set (PR #79
`tests/inference_yaml_pattern_coverage.rs` is the runtime contract).
Adding a new query family here without a matching entry in PR #79's
canonical YAML risks shipping warm-tier "promises" the engine can't
keep — the YAML is checked exact-string at request time by
`find_query_config`.

#### Metric-name conventions per overlay

| Overlay | Backend-side metric name | Why |
|---|---|---|
| `backend-inference.yaml` | `http_requests_total_latency_ms_quantile` (KLL/DDSketch quantile patterns), `http_requests_total` (CMS/CountSketch/HLL Sum/Count/Topk patterns) | Default deploy uses `gateway-aggregate-from-raw.yaml` → DDSketch → `metric_suffix: "_quantile"`. Sum/Count/Topk patterns target the raw counter forwarded unsuffixed by CMS/CountSketch processors (see `sketchcol-agent-{cms,cs}-direct.yaml`). |
| `backend-inference-cms.yaml`, `-cs.yaml` | `http_requests_total` | CMS / CountSketch direct agents preserve the raw metric name (no `metric_suffix`). |
| `backend-inference-hll.yaml` | `http_requests_total_hll`, `http_requests_total_latency_ms_hll` | `sketchcol-agent-hll-direct.yaml` adds `metric_suffix: "_hll"`. |
| `backend-inference-kll.yaml` | `http_requests_total_latency_ms_kll` | `sketchcol-agent-kll-direct.yaml` adds `metric_suffix: "_kll"`. |

#### Open gap: `histogram_quantile(φ, …)` is NOT covered

The engine pattern matcher
(`asap-query-engine/src/engines/simple_engine.rs::controller_patterns`)
includes `quantile_over_time` and the spatial `quantile by (…)` ops
but does **not** include a `histogram_quantile` pattern block.
`histogram_quantile(0.99, sum by (le) (http_requests_total_latency_ms))`
fails the engine matcher before reaching `find_query_config`, even
when an entry of that exact string is present in the YAML.

Two paths forward (in priority order):

1. **Today (this PR's choice for `queries-e2e.json`):** use the
pre-aggregated `quantile_over_time(φ, *_quantile[…])` shape. The
gateway / agent path produces `http_requests_total_latency_ms_quantile`
(DDSketch / KLL backed) which the engine matches and answers.
PROGRESS.md "Single-pipeline multi-sketch + delta + queryable
warm tier (2026-05-01)" verified this path live (`q=0.5 →
19.49`).
2. **Tomorrow (separate engine PR):** extend `controller_patterns`
with a `histogram_quantile` block. Out of scope for E0; tracked
alongside PR #79 follow-ups.

### E0: single-cell smoke (P5–P9 end-to-end against a live stack)

Smallest cell that exercises the whole P1–P9 path. Use this
Expand Down Expand Up @@ -126,6 +179,11 @@ python3 deploy/scripts/plan_transition.py \
--transition-out /tmp/cell-smoke-e0/transition.jsonl \
--sample-out /tmp/cell-smoke-e0/sample.jsonl \
--soak-secs 60 --pre-transition-secs 20 &
# Note: the `histogram_quantile(...)` shape above is INTENTIONALLY
# unmatched by the engine — it's the capability-miss probe used by
# `plan_transition.py` to drive a fresh plan publish. Replay-side
# queries (`queries-e2e.json`) deliberately use shapes that DO match
# (see "Inference dispatch" above).
wait

# 4. Snapshot ground truth from the cold-store volume (the volume
Expand Down
77 changes: 76 additions & 1 deletion deploy/configs/backend-inference-cms.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,27 @@
# Warm-tier dispatch table (CountMinSketch single-sketch overlay).
#
# Subset of `backend-inference.yaml` (which mirrors PR #79's
# `inference_config.yaml`) restricted to the families that route to
# the CountMinSketch accumulator:
#
# * Spatial `sum` / `count` — Statistic::{Sum,Count} no-key
# fallback returns the min-row sum (canonical CMS total-event
# estimator). PROGRESS.md "All-five-sketch runtime e2e
# verification (2026-04-30)" verified
# `sum_over_time(http_requests_total[1m])` → 145735 with ε≈0.0027.
# * `sum_over_time` / `count_over_time` over [1m]/[2m]/[5m] —
# Statistic::Sum / Statistic::Count via
# `additive_frequency` accuracy envelope.
# * `rate` / `increase` — same family (delta arithmetic over Sum).
#
# `topk` / `histogram_quantile` / `quantile_over_time` excluded
# (CMS does not back those statistics — see the
# `compatible_agg_types` table in `capability_matching.rs`).
#
# Metric: `http_requests_total` (no suffix). The CMS direct agent
# overlay (`sketchcol-agent-cms-direct.yaml`) uses
# `metric_name: "http_requests_total"` and preserves the raw name
# end-to-end.
cleanup_policy:
name: "circular_buffer"
metrics:
Expand All @@ -7,11 +31,62 @@ metrics:
- node
- pod
queries:
# ── Spatial Sum / Count / Avg ─────────────────────────────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum(http_requests_total)
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count(http_requests_total)
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: avg(http_requests_total)
# ── sum_over_time / count_over_time × wider ranges ────────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum_over_time(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum(http_requests_total)
query: sum_over_time(http_requests_total[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum_over_time(http_requests_total[5m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total[5m])
# ── rate / increase (delta arithmetic over Sum family) ────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: rate(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: rate(http_requests_total[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: rate(http_requests_total[5m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: increase(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: increase(http_requests_total[5m])
92 changes: 85 additions & 7 deletions deploy/configs/backend-inference-cs.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,26 @@
# Warm-tier dispatch table (CountSketch single-sketch overlay).
#
# Subset of `backend-inference.yaml` (which mirrors PR #79's
# `inference_config.yaml`) restricted to the families that route to
# the CountSketch accumulator:
#
# * `topk(N, …)` — Statistic::Topk; CountSketch's `query_statistic`
# returns the row-mean total when no key is supplied (per PR #79
# wiring; per-key enumeration needs a paired SetAggregator on the
# agent — tracked upstream).
# * Spatial `sum` / `count` — Statistic::{Sum,Count} no-key
# fallback (PROGRESS.md "All-five-sketch runtime e2e verification
# (2026-04-30)" — `sum_over_time(http_requests_total[1m])` →
# 24266 with ε=0.03 `additive_frequency` envelope).
# * `sum_over_time` / `count_over_time` × [1m]/[2m]/[5m].
# * `rate` / `increase` — same Sum family.
#
# `quantile_over_time` / `histogram_quantile` excluded (CountSketch
# does not back rank statistics).
#
# Metric: `http_requests_total` (no suffix). The CountSketch direct
# agent overlay (`sketchcol-agent-cs-direct.yaml`) preserves the raw
# metric name through the processor.
cleanup_policy:
name: "circular_buffer"
metrics:
Expand All @@ -6,17 +29,72 @@ metrics:
- rack
- node
- pod
http_requests_total_latency_ms:
- zone
- rack
- node
- pod
queries:
# ── Spatial Sum / Count ───────────────────────────────────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum(http_requests_total)
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count(http_requests_total)
# ── sum_over_time / count_over_time × wider ranges ────────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum_over_time(http_requests_total[1m])
- aggregations:
- aggregation_id: 2
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum_over_time(http_requests_total[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum_over_time(http_requests_total[5m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total[5m])
# ── rate / increase (delta arithmetic over Sum family) ────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: rate(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: rate(http_requests_total[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: rate(http_requests_total[5m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: increase(http_requests_total[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: increase(http_requests_total[5m])
# ── Top-K (CountSketch heavy-hitter readout) ──────────────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: topk(5, http_requests_total)
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: topk(10, http_requests_total)
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: sum_over_time(http_requests_total_latency_ms[1m])
query: topk(50, http_requests_total)
43 changes: 43 additions & 0 deletions deploy/configs/backend-inference-hll.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,24 @@
# Warm-tier dispatch table (HyperLogLog single-sketch overlay).
#
# Subset of `backend-inference.yaml` (which mirrors PR #79's
# `inference_config.yaml`) restricted to the families that route to
# the HLL accumulator:
#
# * `count(metric)` — Statistic::Count; HLL's `query_statistic`
# accepts both `Cardinality` and `Count` (alias) and returns the
# unique-cardinality estimate. PROGRESS.md "All-five-sketch
# runtime e2e verification (2026-04-30)" verified
# `count(http_requests_total)` → 149.68 distinct, ε=0.008
# `relative_cardinality` envelope.
# * `count_over_time(metric[…])` — temporal Count.
#
# `sum` / `quantile_over_time` / `topk` / `rate` excluded (HLL is a
# cardinality sketch — no value semantics).
#
# Metric: `http_requests_total_hll` / `http_requests_total_latency_ms_hll`
# — the HLL direct agent overlay (`sketchcol-agent-hll-direct.yaml`)
# uses `metric_suffix: "_hll"`, so the backend-side metric name
# is the suffixed form.
cleanup_policy:
name: "circular_buffer"
metrics:
Expand All @@ -12,15 +33,37 @@ metrics:
- node
- pod
queries:
# ── Spatial Cardinality (count over an HLL-backed metric) ─────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count(http_requests_total_hll)
- aggregations:
- aggregation_id: 2
num_aggregates_to_retain: 6
query: count(http_requests_total_latency_ms_hll)
# ── count_over_time × wider ranges ────────────────────────────────────
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total_hll[1m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total_hll[2m])
- aggregations:
- aggregation_id: 1
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total_hll[5m])
- aggregations:
- aggregation_id: 2
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total_latency_ms_hll[1m])
- aggregations:
- aggregation_id: 2
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total_latency_ms_hll[2m])
- aggregations:
- aggregation_id: 2
num_aggregates_to_retain: 6
query: count_over_time(http_requests_total_latency_ms_hll[5m])
Loading