Skip to content

MVP demo: test-first validation of ASAPCollector + ASAPQuery-backend #46

Description

@zzylol

MVP demo: test-first validation of ASAPCollector + ASAPQuery-backend

Goal

Demonstrate, with one reproducible end-to-end test harness, that ASAPCollector and ASAPQuery-backend form a correct and useful sketch-based metrics pipeline:

  1. the control plane selects what the OTel collector computes and transmits;
  2. ASAPQuery-backend answers supported queries from those sketches with bounded error and fresh results;
  3. sketch-backed queries are faster than exact queries over Prometheus or VictoriaMetrics; and
  4. the edge-only and end-to-end resource costs are lower than the exact raw-metrics baseline.

The test harness and its pass/fail contract must be written before optimizing individual components. A fresh run must produce all evidence used to declare the MVP complete.

MVP architecture under test

Identical generated/replayed metrics
  ├─ ASAP arm: OTLP → ASAPCollector OTel processors → sketches/full or delta
  │            → ASAPQuery-backend → approximate PromQL result
  └─ Exact arm: raw metrics → Prometheus or VictoriaMetrics
               → exact PromQL result

Query workload → control plane decision → collector configuration

Both arms must receive the same samples, timestamps, labels, query workload, and test duration. Measurements must use steady-state intervals after a documented warm-up. Prometheus or VictoriaMetrics may be used as the exact system, but the selected baseline and its configuration must be recorded in the report.

Deliverable 1: end-to-end test harness (implement first)

Provide one command that starts both arms, generates or replays the workload, submits the query workload to the control plane, runs correctness and benchmark queries, collects resource measurements, and emits a machine-readable result plus a human-readable MVP_REPORT.md.

The harness must:

  • pin and report the commit SHA and configuration of ASAPCollector, ASAPQuery-backend, the load generator, and the exact backend;
  • run identical input data through the ASAP and exact arms;
  • cover representative supported query classes:
    • aggregation within a time window for each series;
    • aggregation across label groups at a timestamp;
    • aggregation across both a selected time window and label groups;
  • exercise every control-plane transmission decision used by the MVP: raw/pass-through when selected, full-sketch transmission, and delta-sketch transmission for sketch families that support delta encoding;
  • prove that the collector applied the issued decision, rather than only checking the controller response;
  • query both systems over the same logical time range and compare aligned result series;
  • measure query latency over repeated runs and report at least p50 and p95, with warm-up queries excluded;
  • measure result freshness from sample generation timestamp until that sample/window is queryable;
  • measure CPU time, peak/steady-state memory, and bytes transmitted separately for the collector and backend processes;
  • measure exact-baseline ingestion, storage, and query resources over the same run;
  • preserve raw observations, configs, logs, query responses, and derived measurements as run artifacts;
  • fail closed: missing, stale, empty, or incomparable measurements are failures, not PASS or silently skipped checks.

The report must include the workload shape (sample rate, series cardinality, label cardinality, duration, query mix, and sketch parameters) and enough commands/configuration to reproduce the run.

Deliverable 2: MVP acceptance criteria

All criteria below must pass in the same reproducible end-to-end run. Exact numerical thresholds that are workload-dependent must be declared in the checked-in test configuration before the measurement run; they must not be selected after observing results.

1. Functional correctness

  • The control plane converts the submitted query workload into an explicit collection plan containing the selected sketch family, parameters, aggregation/window placement, and transmission mode.
  • The ASAPCollector OTel processor receives and applies that plan, computes the selected sketches, and sends either full sketches or deltas as directed.
  • ASAPQuery-backend ingests the emitted payloads and returns valid Prometheus-compatible results for every supported MVP query.
  • Full and delta transmission produce equivalent query semantics within the same configured accuracy bound.
  • Unsupported queries are explicitly rejected or routed to the configured exact fallback; they must never return a plausible but incorrect sketch result.
  • Component failures, dropped payloads, and query errors are surfaced by the harness.

2. Query accuracy

  • The exact Prometheus/VictoriaMetrics result is the ground truth for the identical input stream and query interval.
  • Every returned series is aligned by labels and timestamp before comparison; missing or extra series count as errors.
  • Each query declares its accuracy metric and SLA (for example relative/absolute error for numeric aggregates, rank/quantile error, or precision/recall for set and heavy-hitter queries).
  • The report includes per-query median, p95, p99, and worst-case error, plus the fraction of answers within the configured SLA.
  • Pass condition: all supported MVP queries satisfy their predeclared accuracy SLA. Exact operators must match the exact baseline within numeric tolerance; approximate operators must satisfy their configured sketch error bound.

3. Query freshness

  • Freshness is measured end to end, from the source sample timestamp to the first successful ASAPQuery-backend query that includes the corresponding data/window.
  • Report p50, p95, p99, and maximum freshness lag for each query class and for both full- and delta-sketch modes where applicable.
  • Pass condition: every supported query meets its predeclared freshness SLA, and no result used for accuracy or latency scoring is stale or from an earlier run.

4. Query performance

  • Compare ASAPQuery-backend sketch-query latency with exact query latency from the selected Prometheus/VictoriaMetrics baseline using the same query semantics, data interval, cardinality, client, concurrency, and number of repetitions.
  • Report cold and warm results separately when both are relevant; do not compare a warm ASAP query with a cold exact query.
  • Pass condition: ASAPQuery-backend has lower steady-state p50 and p95 latency than the exact baseline for every query class claimed as accelerated. Report the absolute latency and speedup ratio; statistical dispersion must also be included.

5. Collector cost reduction

Compare the ASAPCollector arm with an OTel/raw-forwarding collector that transmits every observed raw metric to the same exact backend under the same input load.

  • Measure collector CPU time, peak and steady-state memory, and transmitted bytes.
  • Report each resource separately; do not hide a regression inside an unweighted sum of unlike units.
  • Define a checked-in cost model that converts CPU, memory occupancy, and network bytes into one normalized cost using fixed weights.
  • Pass condition: ASAPCollector's normalized total cost is lower than the raw-forwarding collector baseline, and any individual-resource regression is disclosed and remains below its predeclared guardrail.

6. End-to-end cost reduction

Compare the complete systems over the same ingest-and-query workload:

  • ASAP: ASAPCollector compute/memory/network + ASAPQuery-backend ingest/query compute/memory + sketch storage;
  • Exact baseline: raw collector compute/memory/network + Prometheus/VictoriaMetrics ingest/query compute/memory + raw-data storage.

Use measured resource quantities and the same checked-in cost weights for both systems. Include ingestion, transmission, storage, and query execution; excluding an unfavorable stage invalidates the comparison.

  • Pass condition: the normalized total cost of ASAPCollector + ASAPQuery-backend is lower than the full raw-data exact baseline while the functional, accuracy, freshness, and latency gates above remain satisfied.

Required output summary

MVP_REPORT.md must end with a table like this, backed by links to run artifacts:

Category Metric ASAP Exact baseline Required result Verdict
Functional correctness passed scenarios / total all
Query accuracy per-query error/SLA within SLA
Query freshness p95 lag within SLA
Query performance p50/p95 latency ASAP lower
Collector cost normalized cost ASAP lower
End-to-end cost normalized cost ASAP lower

The overall MVP verdict is PASS only when every category passes. UNKNOWN, missing data, skipped scenarios, or results from a previous run make the overall verdict FAIL.

Definition of done

  • The end-to-end harness and checked-in acceptance configuration exist before further benchmark-specific optimization.
  • One documented command runs the paired ASAP and exact-baseline experiment from a clean environment.
  • Control-plane decisions and the collector's applied configuration are captured as artifacts.
  • Functional correctness passes for the three MVP query classes and all claimed transmission modes.
  • Accuracy SLAs pass against exact, aligned results.
  • Freshness SLAs pass with no stale-run artifacts.
  • ASAPQuery-backend p50 and p95 query latency are lower than exact computation for every claimed accelerated query class.
  • ASAPCollector normalized cost is lower than raw forwarding, with per-resource guardrails reported.
  • Total ASAPCollector + ASAPQuery-backend normalized cost is lower than the full exact pipeline cost.
  • A fresh MVP_REPORT.md, machine-readable results, raw measurements, logs, and configs are retained for review.

Non-goals for this MVP

  • Proving performance for PromQL operators that are not claimed as sketch-accelerated.
  • Choosing thresholds after benchmark results are known.
  • Treating projected or analytically estimated savings as substitutes for end-to-end measurements.
  • Requiring a specific headline ingest rate (for example, 1M samples/s) unless that scale is declared in the test configuration as a separate scalability experiment.
  • Comparing against Serf or validating archive/cold-store fallback unless those are added as separately scoped acceptance criteria.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions