Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
-
Updated
Sep 28, 2026 - Python
Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
Build trustworthy pytest-benchmark matrices, collect repeated runs, and detect performance regressions
LLM agent that turns two Android Perfetto traces and a git range into a cited diagnosis: the regressed metric, the culprit commit, and the trace rows that prove it. Every claim is re-verified by code; measured on planted regressions.
Evidence-backed performance regression guardrails for pull requests, with protected verification and bounded Codex repair workflows.
Profile a pytest suite, script, or callable and comment a CPU + memory hotspot diff on every pull request.
Compare Rust benchmark results across pull-request revisions
To associate your repository with the performance-regression topic, visit your repo's landing page and select "manage topics."