Skip to content

feat(asap-tools): label query sets and skew fits for cluster traces - #746

Draft
zzylol wants to merge 9 commits into
mainfrom
feat/trace-label-skew
Draft

zzylol wants to merge 9 commits into
mainfrom
feat/trace-label-skew

Conversation

@zzylol

@zzylol zzylol commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds asap-tools/dataset-analysis/, which covers two tasks:

  1. Label query sets: PromQL queries, with their exact group-by labels, for three traces:
    • Google ClusterData 2011
    • Alibaba microservices v2022
    • Datadog BOOM
  2. Skew fits: for each query, a lower bound, a maximum-likelihood fit and an upper bound of the skew parameter.
    • Queries that aggregate over keys (the group-by labels) get a discrete Zipf θ, weighted by row count and by value sum.
    • Queries that aggregate over values get a power-law tail α. It is fitted with powerlaw, the method of Clauset, Shalizi & Newman, and compared against lognormal and exponential fits.
    • lower/upper are the min/max of per-window MLEs; mle is the fit on all the data.
    • Windows are 5 min for Google, 1 min for Alibaba, and 20 equal chunks per series for BOOM.

results/skew_summary.csv (59 rows) is the handoff to the sketch-bench saturation study.

Conclusions (tasks 1–2: how skewed is real monitoring data)

  1. Real group-by keys are skewed, with θ ≈ 1.0–1.5. This holds for Google user / job / priority and Alibaba service / msname / um / dm. Only per-machine or per-node keys, where each key is one series, are near-uniform (θ ≈ 0–0.3).
  2. The worst case is much worse than the average, and longer samples expose it. For Google by (user), the lower bound on θ is 1.22 on a 1.4 h sample but 1.10 on a 7-day sample with 5-minute evaluations. Fitting on a short sample, or on average behaviour, undersizes the sketch.
  3. Instant and range queries see different distributions. An instant query counts series per group, while a range query counts samples per group. For MSRTMCR by (msname), θ is 0.83 for the instant query and 1.44 for 5m/1h ranges, and the instant query needs a 4× larger CMS. Fits must follow the query's semantics.
  4. The number of distinct keys K per window depends on the time scale. CallGraph by (service) has about 40K keys per minute and 2.77M over 6 h. Google by (user) stays at about 500–700 keys.
  5. Most value tails are lognormal or light; few are truly heavy. The clearly heavy ones are call latency rt (α ≈ 2) and a few BOOM series. Power law vs lognormal is often undecidable, so we report a tail class (light / power_law / lognormal / heavy_inconclusive) rather than a pass/fail power-law test.

Data

fetch_data.sh downloads a fixed sample:

  • Google: task_usage and task_events part-00000, plus all of job_events.
  • Alibaba v2022: NodeMetrics_0, MSMetrics_0, and 10 consecutive CallGraph and MCRRT shards (30 min). These are read straight from tar.gz.
  • BOOM: 20 sampled series plus the taxonomy.

A full run takes about 4.5 min with 48 workers.

Key θ (lower / mle / upper), count-weighted; value-weighted in brackets

Query θ
Google sum by (user) (cpu_rate) 1.22 / 1.24 / 1.28 [1.25 / 1.30 / 1.42]
Google sum by (user, priority) 1.23 / 1.25 / 1.28 [1.27 / 1.32 / 1.44]
Google sum by (priority) 1.03 / 1.10 / 1.17 [1.44 / 1.61 / 1.75]
Google sum by (job_id) 1.11 / 1.10 / 1.14
Google sum by (scheduling_class) 0.41 / 0.58 / 0.76
Google sum by (machine_id) (uniform baseline) 0.23 / 0.21 / 0.25
Alibaba MSRTMCR by (msname) 1.34 / 1.54 / 1.90 [1.03 / 1.11 / 1.31]
Alibaba MSRTMCR by (nodeid) 1.07 / 1.13 / 1.23
CallGraph count by (rpctype) 1.26 / 1.32 / 1.39
CallGraph count by (um) / by (dm) 1.11 / 1.15 / 1.13, 1.10 / 1.13 / 1.11
CallGraph count by (service, um, dm) 0.91 / 0.98 / 0.93
MSMetrics sum by (msname) 0.83 (the same in every window)
NodeMetrics by (nodeid) (uniform baseline) ≈0

Known issues (why this is a draft)

  • mle can fall outside [lower, upper]. For most CallGraph queries, the fit on all 30 minutes is steeper than the fit on any single minute: pooling adds rare keys to the tail and more mass to the heavy keys. So mle is not bracketed by the window bounds. The definition of the bounds needs a decision.
  • Value α is not ready to use as-is.
    • Fits use up to 5,000 sampled values, and the tail must keep at least 100 values, so α describes roughly the top ≥2% of values.
    • The Google fits lose to lognormal.
    • For BOOM, power_law_ok is a majority vote over variates while R and p are medians. As a result ds-671-10S and ds-1840-D show ok=True with absurd α.
  • powerlaw is pinned to 1.5. Version 2.0 caps α at 3 by default and falls back to a slow numerical fit.
  • BOOM: the tags were removed and each variate was z-scored. So there is no key θ for BOOM, and α is fitted per variate on x - min(x).
  • The um/dm labels include the placeholder values UNKNOWN and USER, which are counted as ordinary keys.

Also changed

.github/workflows/python.yml gets a test-dataset-analysis job (black, isort, flake8, mypy, unittest), path-filtered to this directory.

Validation

  • black, isort, flake8, mypy and shellcheck pass.
  • 18 unit tests pass. They cover Zipf θ recovery within ±0.05, Pareto α recovery within ±5%, the bounds logic, and failure paths.

Citations

  • Google ClusterData 2011 traces (2025)
  • Luo et al., "The Power of Prediction: Microservice Auto Scaling via Workload Learning", SoCC 2022
  • Cohen et al., "This Time is Different: An Observability Perspective on Time Series Foundation Models", arXiv 2505.14766

🤖 Generated with Claude Code


Update: query-range-aligned fits, longer sample, worst-case parameters

The sections above describe the first version. Since then:

  • Longer sample:
    • Google: 7 days (120 task_usage and 120 task_events parts).
    • Alibaba v2022: CallGraph and MCRRT for 6 hours (120 shards each); MSMetrics and NodeMetrics for 1 day.
    • About 63 GB total, downloaded by fetch_data.sh.
  • Query semantics. Each query has a range field:
    • Instant (metric tables): at every evaluation step, take each series' latest sample within the 5-minute lookback, then group by the query labels.
    • Range (5m, 1h; CallGraph 1m, 5m, 1h): all samples in (t − S, t], sliding one step at a time. Steps are 1 min for Alibaba and 5 min for Google.
    • Bounds: lower/upper are the min/max across evaluation times; mle is the fit on all data.
  • Tail classes replace the pass/fail power-law flag. Each value query gets one of light, power_law, lognormal or heavy_inconclusive; whether something is a power law or a lognormal often can't be decided (Clauset et al. 2009).
  • Columns for sketch-bench#130's recommend_config.py:
    • Worst-case parameters: worst_theta_cms, worst_K, min_N, max_N, worst_alpha_rank, worst_alpha_memory.
    • Per-query accuracy targets: target_*.
    • Per-evaluation K_win_* and rows_win_*.
  • Notable result. For MSRTMCR by (msname), the instant θ is 0.83 but the 5m/1h θ is 1.44: an instant query counts series per key, a range query counts rows per key. So an instant query needs a 4× larger CMS (3×16384 vs 3×4096) for the same accuracy target.
  • Run cost. A full run takes about 2 h with 24 workers and peaks at 69 GB of memory. 42 unit tests; black, isort, flake8, mypy and shellcheck are clean.

zzylol and others added 9 commits September 29, 2026 18:21
Add asap-tools/dataset-analysis: query sets over Google 2011, Alibaba
v2022 and BOOM traces, a fetch script, and fit_skew.py, which fits a
discrete Zipf theta per key query and a power-law alpha per value query,
with lower/upper bounds from per-window fits. Add a CI job running its
lint, type check and unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Output of fit_skew.py over the data fetched by fetch_data.sh.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- lower/upper now span the per-window fits and the pooled fit, so
  lower <= mle <= upper.
- Bounds are computed at several window lengths (window_lengths_s per
  dataset) by merging finest-window aggregates; results gain window_len_s.
- Value fits use up to 100k samples and pick xmin over a 50-quantile grid
  (p50..p99.9, at least 100 tail values); losing fits report best_alt.
- BOOM series report ok_frac and alpha over passing variates only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
power_law_ok now needs the power law to significantly beat both the
lognormal and the exponential. Otherwise best_alt names the significantly
better alternative or "inconclusive". BOOM ok_frac uses the same rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace power_law_ok and best_alt with tail_class (light, power_law,
lognormal, heavy_inconclusive). BOOM series report the majority class,
ok_frac as the share of non-light variates, and alpha and R/p medians
over the non-light variates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
K and rows were whole-sample totals repeated on every window length.
Rename them to K_total/rows_total and add K_win_{min,median,max} and
rows_win_{min,median,max} from the existing per-window aggregates
(value queries get rows_win_* only), so the saturation study can read N
and K at the sketch window and the query lookback. Write the summary
with enough digits to keep row counts exact.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fetch and analyze a longer sample: Google task_usage/task_events parts
0-119 (about 7 days), Alibaba CallGraph/MCRRTUpdate shards 0-119 (6 h),
MSMetricsUpdate 0-47 and NodeMetricsUpdate 0-1 (1 day). Drop the
30-minute clip on Alibaba tables and let a table override the dataset's
window_lengths_s, so window lengths reach 6 h (CallGraph, MSRTMCR), 1 day
(MSMetrics, NodeMetrics) and 7 days (Google). Joins are loaded once per
table instead of once per file, so the task_events join covers all 120
parts. Regenerate results/skew_summary.csv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… queries

Replace disjoint window lengths with Prometheus-style evaluation: every
query is evaluated at each step of its table (step_s, the sampling
period), in an instant form (latest sample per series, 5-minute
lookback) and/or range forms (sliding window over the last S seconds).
lower/upper are the min/max over evaluation times; mle pools all data.

- Tables declare step_s and series_key; queries declare promql,
  promql_range and range (instant, 5m, 1h; CallGraph events: 1m, 5m, 1h).
  Google task_usage is stamped at end_time with a 5-minute step.
- Range key sums use a banded sparse matrix over per-step key
  aggregates; value queries merge per-step uniform samples.
- Instant aggregates are computed per file, with series boundary
  samples resolved across files; interleaved files fail loudly.
- Add worst-case columns for the sketch-bench study (worst_theta_cms,
  worst_K, min_N, max_N, worst_alpha_rank, worst_alpha_memory) and
  per-query accuracy targets; cms_weight picks the counter-sketch weight.
- Evaluations with no rows (trace gaps) are skipped for key and value
  queries alike.
- Regenerate results/skew_summary.csv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rank-frequency and CCDF plots for every query and range from the current
results, with a short guide to the file names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant