Repository navigation
Make CI benchmark comparison robust (median warm metrics + filesystem warmup + tests) - #511
Merged
Merged
Conversation
Up to standards ✅🟢 Issues
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #511 +/- ##
=======================================
Coverage 83.36% 83.36%
=======================================
Files 86 86
Lines 18997 18997
Branches 18997 18997
=======================================
Hits 15836 15836
Misses 2847 2847
Partials 314 314 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Contributor
📊 Benchmark ComparisonCurrent:
Compared Liquid vs DataFusionDefault on the same runner |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Description
.github/compare_benchmarks.pyby importingstatisticsand updatingget_warm_metricsto return median values.--thresholdconsistently used when highlighting regressions and computing the warm-time summary informat_change_percentageand the report generation..github/test_compare_benchmarks.pythat verifies median warm metrics and threshold-based highlighting..github/workflows/ci.ymlto build the releasein_processbinary once, run a small unmeasured DataFusion warmup to populate the filesystem page cache, and invoketarget/release/in_processfor measured runs to remove order-dependent bias.Testing
python3 .github/test_compare_benchmarks.py, and they passed (2 tests, OK).ruffandpython -m py_compile, both passed.cargo check -p liquid-cache-benchmarksfor the benchmark crate succeeded.cargo checkencountered a pre-existing environment-specific failure due to a missing generated asset (dev/dev-tools/assets/tailwind.css) that is unrelated to these benchmark changes.Codex Task