Skip to content

Make CI benchmark comparison robust (median warm metrics + filesystem warmup + tests) - #511

Merged
XiangpengHao merged 1 commit into
mainfrom
codex/analyze-and-fix-ci-benchmark-results
Sep 2, 2026
Merged

XiangpengHao merged 1 commit into
mainfrom
codex/analyze-and-fix-ci-benchmark-results

Conversation

@XiangpengHao

Copy link
Copy Markdown
Collaborator

Motivation

  • The CI benchmark comparison showed large, order-dependent slowdowns because LiquidCache runs first and populates the OS page cache, making the baseline look artificially faster.
  • Single outlier warm iterations from CI pauses can make the mean-driven warm metric noisy and report false regressions.
  • Make minimal, targeted changes to the comparison and CI workflow so results reflect real performance differences.

Description

  • Replace mean with median for warm-iteration metrics in .github/compare_benchmarks.py by importing statistics and updating get_warm_metrics to return median values.
  • Make the configured --threshold consistently used when highlighting regressions and computing the warm-time summary in format_change_percentage and the report generation.
  • Add a lightweight unit test file .github/test_compare_benchmarks.py that verifies median warm metrics and threshold-based highlighting.
  • Update CI in .github/workflows/ci.yml to build the release in_process binary once, run a small unmeasured DataFusion warmup to populate the filesystem page cache, and invoke target/release/in_process for measured runs to remove order-dependent bias.

Testing

  • Ran unit tests python3 .github/test_compare_benchmarks.py, and they passed (2 tests, OK).
  • Linted and checked Python files with ruff and python -m py_compile, both passed.
  • Verified cargo check -p liquid-cache-benchmarks for the benchmark crate succeeded.
  • A full workspace cargo check encountered a pre-existing environment-specific failure due to a missing generated asset (dev/dev-tools/assets/tailwind.css) that is unrelated to these benchmark changes.

Codex Task

@codacy-production

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.

@codecov

codecov Bot commented Sep 1, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 83.36%. Comparing base (1909d08) to head (8015b34).

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #511   +/-   ##
=======================================
  Coverage   83.36%   83.36%           
=======================================
  Files          86       86           
  Lines       18997    18997           
  Branches    18997    18997           
=======================================
  Hits        15836    15836           
  Misses       2847     2847           
  Partials      314      314           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

📊 Benchmark Comparison

Current: d61cd179 (Liquid) vs Baseline: d61cd179 (DataFusionDefault)

Query Cold Time Δ Warm Time Δ CPU Time Δ
Q1 3.0ms (3.0ms) +0.0% 0.000ms (0.000ms) +0.0% 0.000ms (0.000ms) +0.0%
Q2 14.0ms (6.0ms) +133.3% 6.0ms (6.5ms) -7.7% 6.0ms (7.5ms) -20.0%
Q3 19.0ms (12.0ms) +58.3% 7.0ms (12.0ms) -41.7% 2.5ms (23.0ms) -89.1%
Q4 20.0ms (18.0ms) +11.1% 5.5ms (11.5ms) -52.2% 2.0ms (24.5ms) -91.8%
Q5 80.0ms (54.0ms) +48.1% 92.5ms (53.5ms) +72.9% 4.5ms (26.0ms) -82.7%
Q6 187.0ms (109.0ms) +71.6% 100.0ms (108.5ms) -7.8% 35.5ms (78.0ms) -54.5%
Q7 3.0ms (1.0ms) +200.0% 0.500ms (1.0ms) -50.0% 0.000ms (0.000ms) +0.0%
Q8 12.0ms (10.0ms) +20.0% 8.0ms (5.5ms) +45.5% 9.5ms (6.5ms) +46.2%
Q9 125.0ms (86.0ms) +45.3% 113.5ms (84.0ms) +35.1% 4.0ms (44.0ms) -90.9%
Q10 117.0ms (87.0ms) +34.5% 93.0ms (94.5ms) -1.6% 5.0ms (63.0ms) -92.1%
Q11 48.0ms (25.0ms) +92.0% 21.5ms (28.0ms) -23.2% 39.5ms (37.5ms) +5.3%
Q12 49.0ms (29.0ms) +69.0% 24.5ms (28.0ms) -12.5% 40.0ms (43.5ms) -8.0%
Q13 204.0ms (108.0ms) +88.9% 115.0ms (107.0ms) +7.5% 69.0ms (79.5ms) -13.2%
Q14 363.0ms (159.0ms) +128.3% 157.0ms (141.0ms) +11.3% 113.5ms (109.0ms) +4.1%
Q15 227.0ms (107.0ms) +112.1% 116.5ms (105.0ms) +11.0% 82.0ms (95.0ms) -13.7%
Q16 147.0ms (114.0ms) +28.9% 129.5ms (101.5ms) +27.6% 6.0ms (25.5ms) -76.5%
Q17 443.0ms (207.0ms) +114.0% 243.0ms (214.5ms) +13.3% 88.0ms (103.5ms) -15.0%
Q18 440.0ms (211.0ms) +108.5% 239.0ms (207.0ms) +15.5% 87.0ms (103.5ms) -15.9%
Q19 669.0ms (386.0ms) +73.3% 385.5ms (432.5ms) -10.9% 110.5ms (148.5ms) -25.6%
Q20 14.0ms (11.0ms) +27.3% 3.0ms (12.5ms) -76.0% 7.0ms (24.0ms) -70.8%
Q21 949.0ms (173.0ms) +448.6% 262.0ms (173.0ms) +51.4% 510.5ms (271.0ms) +88.4%
Q22 1.11s (161.0ms) +591.3% 389.5ms (167.0ms) +133.2% 179.5ms (342.5ms) -47.6%
Q23 2.32s (460.0ms) +404.3% 902.5ms (458.5ms) +96.8% 532.5ms (733.5ms) -27.4%
Q24 23.10s (907.0ms) +2446.6% 805.5ms (945.0ms) -14.8% 712.0ms (2.62s) -72.8%
Q25 192.0ms (70.0ms) +174.3% 15.0ms (57.0ms) -73.7% 36.5ms (120.0ms) -69.6%
Q26 88.0ms (44.0ms) +100.0% 21.0ms (46.5ms) -54.8% 53.0ms (84.0ms) -36.9%
Q27 180.0ms (59.0ms) +205.1% 26.5ms (57.5ms) -53.9% 77.0ms (123.5ms) -37.7%
Q28 1.03s (214.0ms) +379.9% 296.0ms (212.5ms) +39.3% 375.0ms (274.0ms) +36.9%
Q29 1.78s (1.06s) +67.5% 1.11s (997.5ms) +11.7% 610.5ms (354.0ms) +72.5%
Q30 29.0ms (30.0ms) -3.3% 25.5ms (30.0ms) -15.0% 7.0ms (21.5ms) -67.4%
Q31 326.0ms (104.0ms) +213.5% 72.0ms (113.0ms) -36.3% 43.5ms (152.0ms) -71.4%
Q32 608.0ms (109.0ms) +457.8% 99.0ms (107.0ms) -7.5% 58.0ms (154.5ms) -62.5%
Q33 357.0ms (318.0ms) +12.3% 326.0ms (320.0ms) +1.9% 7.0ms (73.5ms) -90.5%
Q34 1.15s (432.0ms) +165.7% 535.0ms (418.0ms) +28.0% 361.0ms (281.0ms) +28.5%
Q35 1.19s (397.0ms) +199.7% 486.0ms (424.0ms) +14.6% 372.0ms (275.5ms) +35.0%
Q36 105.0ms (100.0ms) +5.0% 91.5ms (106.0ms) -13.7% 3.5ms (27.0ms) -87.0%
Q37 315.0ms (106.0ms) +197.2% 91.5ms (103.0ms) -11.2% 45.0ms (71.0ms) -36.6%
Q38 69.0ms (49.0ms) +40.8% 33.5ms (44.5ms) -24.7% 20.5ms (23.5ms) -12.8%
Q39 258.0ms (45.0ms) +473.3% 13.0ms (52.5ms) -75.2% 11.0ms (74.0ms) -85.1%
Q40 829.0ms (201.0ms) +312.4% 227.5ms (205.5ms) +10.7% 92.5ms (134.5ms) -31.2%
Q41 23.0ms (19.0ms) +21.1% 11.0ms (19.0ms) -42.1% 7.0ms (16.5ms) -57.6%
Q42 24.0ms (17.0ms) +41.2% 10.5ms (17.5ms) -40.0% 6.5ms (14.5ms) -55.2%
Q43 21.0ms (15.0ms) +40.0% 11.0ms (15.0ms) -26.7% 6.0ms (10.0ms) -40.0%

⚠️ LiquidCache is slower on 10 queries (warm)

  • Q22: warm +133.2% (389.5ms vs 167.0ms)
  • Q23: warm +96.8% (902.5ms vs 458.5ms)
  • Q5: warm +72.9% (92.5ms vs 53.5ms)
  • Q21: warm +51.4% (262.0ms vs 173.0ms)
  • Q8: warm +45.5% (8.0ms vs 5.5ms)
  • Q28: warm +39.3% (296.0ms vs 212.5ms)
  • Q9: warm +35.1% (113.5ms vs 84.0ms)
  • Q34: warm +28.0% (535.0ms vs 418.0ms)
  • Q16: warm +27.6% (129.5ms vs 101.5ms)
  • Q18: warm +15.5% (239.0ms vs 207.0ms)

Compared Liquid vs DataFusionDefault on the same runner
Regressions: warm-time increases of at least 15%. Cold Time: first iteration; Warm Time: median of remaining iterations.

@XiangpengHao
XiangpengHao merged commit 4761d52 into main Sep 2, 2026
15 checks passed
@XiangpengHao
XiangpengHao deleted the codex/analyze-and-fix-ci-benchmark-results branch September 2, 2026 01:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant