Conversation
Add asap-tools/dataset-analysis: query sets over Google 2011, Alibaba v2022 and BOOM traces, a fetch script, and fit_skew.py, which fits a discrete Zipf theta per key query and a power-law alpha per value query, with lower/upper bounds from per-window fits. Add a CI job running its lint, type check and unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Output of fit_skew.py over the data fetched by fetch_data.sh. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- lower/upper now span the per-window fits and the pooled fit, so lower <= mle <= upper. - Bounds are computed at several window lengths (window_lengths_s per dataset) by merging finest-window aggregates; results gain window_len_s. - Value fits use up to 100k samples and pick xmin over a 50-quantile grid (p50..p99.9, at least 100 tail values); losing fits report best_alt. - BOOM series report ok_frac and alpha over passing variates only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
power_law_ok now needs the power law to significantly beat both the lognormal and the exponential. Otherwise best_alt names the significantly better alternative or "inconclusive". BOOM ok_frac uses the same rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace power_law_ok and best_alt with tail_class (light, power_law, lognormal, heavy_inconclusive). BOOM series report the majority class, ok_frac as the share of non-light variates, and alpha and R/p medians over the non-light variates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
K and rows were whole-sample totals repeated on every window length.
Rename them to K_total/rows_total and add K_win_{min,median,max} and
rows_win_{min,median,max} from the existing per-window aggregates
(value queries get rows_win_* only), so the saturation study can read N
and K at the sketch window and the query lookback. Write the summary
with enough digits to keep row counts exact.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fetch and analyze a longer sample: Google task_usage/task_events parts 0-119 (about 7 days), Alibaba CallGraph/MCRRTUpdate shards 0-119 (6 h), MSMetricsUpdate 0-47 and NodeMetricsUpdate 0-1 (1 day). Drop the 30-minute clip on Alibaba tables and let a table override the dataset's window_lengths_s, so window lengths reach 6 h (CallGraph, MSRTMCR), 1 day (MSMetrics, NodeMetrics) and 7 days (Google). Joins are loaded once per table instead of once per file, so the task_events join covers all 120 parts. Regenerate results/skew_summary.csv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… queries Replace disjoint window lengths with Prometheus-style evaluation: every query is evaluated at each step of its table (step_s, the sampling period), in an instant form (latest sample per series, 5-minute lookback) and/or range forms (sliding window over the last S seconds). lower/upper are the min/max over evaluation times; mle pools all data. - Tables declare step_s and series_key; queries declare promql, promql_range and range (instant, 5m, 1h; CallGraph events: 1m, 5m, 1h). Google task_usage is stamped at end_time with a 5-minute step. - Range key sums use a banded sparse matrix over per-step key aggregates; value queries merge per-step uniform samples. - Instant aggregates are computed per file, with series boundary samples resolved across files; interleaved files fail loudly. - Add worst-case columns for the sketch-bench study (worst_theta_cms, worst_K, min_N, max_N, worst_alpha_rank, worst_alpha_memory) and per-query accuracy targets; cms_weight picks the counter-sketch weight. - Evaluations with no rows (trace gaps) are skipped for key and value queries alike. - Regenerate results/skew_summary.csv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rank-frequency and CCDF plots for every query and range from the current results, with a short guide to the file names. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
asap-tools/dataset-analysis/, which covers two tasks:powerlaw, the method of Clauset, Shalizi & Newman, and compared against lognormal and exponential fits.lower/upperare the min/max of per-window MLEs;mleis the fit on all the data.results/skew_summary.csv(59 rows) is the handoff to the sketch-bench saturation study.Conclusions (tasks 1–2: how skewed is real monitoring data)
by (user), the lower bound on θ is 1.22 on a 1.4 h sample but 1.10 on a 7-day sample with 5-minute evaluations. Fitting on a short sample, or on average behaviour, undersizes the sketch.by (msname), θ is 0.83 for the instant query and 1.44 for 5m/1h ranges, and the instant query needs a 4× larger CMS. Fits must follow the query's semantics.by (service)has about 40K keys per minute and 2.77M over 6 h. Googleby (user)stays at about 500–700 keys.rt(α ≈ 2) and a few BOOM series. Power law vs lognormal is often undecidable, so we report a tail class (light / power_law / lognormal / heavy_inconclusive) rather than a pass/fail power-law test.Data
fetch_data.shdownloads a fixed sample:task_usageandtask_eventspart-00000, plus all ofjob_events.A full run takes about 4.5 min with 48 workers.
Key θ (lower / mle / upper), count-weighted; value-weighted in brackets
sum by (user) (cpu_rate)sum by (user, priority)sum by (priority)sum by (job_id)sum by (scheduling_class)sum by (machine_id)(uniform baseline)by (msname)by (nodeid)count by (rpctype)count by (um)/by (dm)count by (service, um, dm)sum by (msname)by (nodeid)(uniform baseline)Known issues (why this is a draft)
mlecan fall outside[lower, upper]. For most CallGraph queries, the fit on all 30 minutes is steeper than the fit on any single minute: pooling adds rare keys to the tail and more mass to the heavy keys. Somleis not bracketed by the window bounds. The definition of the bounds needs a decision.power_law_okis a majority vote over variates while R and p are medians. As a result ds-671-10S and ds-1840-D showok=Truewith absurd α.powerlawis pinned to 1.5. Version 2.0 caps α at 3 by default and falls back to a slow numerical fit.x - min(x).um/dmlabels include the placeholder valuesUNKNOWNandUSER, which are counted as ordinary keys.Also changed
.github/workflows/python.ymlgets atest-dataset-analysisjob (black, isort, flake8, mypy, unittest), path-filtered to this directory.Validation
Citations
🤖 Generated with Claude Code
Update: query-range-aligned fits, longer sample, worst-case parameters
The sections above describe the first version. Since then:
task_usageand 120task_eventsparts).fetch_data.sh.rangefield:light,power_law,lognormalorheavy_inconclusive; whether something is a power law or a lognormal often can't be decided (Clauset et al. 2009).recommend_config.py:worst_theta_cms,worst_K,min_N,max_N,worst_alpha_rank,worst_alpha_memory.target_*.K_win_*androws_win_*.by (msname), the instant θ is 0.83 but the 5m/1h θ is 1.44: an instant query counts series per key, a range query counts rows per key. So an instant query needs a 4× larger CMS (3×16384 vs 3×4096) for the same accuracy target.