Skip to content

Eval caches include the GPU count and type - #455

Merged
renmengye merged 2 commits into
mainfrom
fix/eval-cache-gpu-count
Oct 2, 2026
Merged

renmengye merged 2 commits into
mainfrom
fix/eval-cache-gpu-count

Conversation

@renmengye

Copy link
Copy Markdown
Member

The dispatched-eval determinant (DispatchedMeasurer._det) omitted the GPU count, although the baseline cache checks it and its docstring said the determinant covers it. A benchmark whose gpus changed, with the same image, command, tree and environment, could reuse a result measured with a different GPU count. The pre-budget lookup could also discount the gate charge on that stale entry. Found while reviewing the accelerators design (#454, phase 0).

  • The determinant now includes the GPU count and the resolved GPU type from the lane, in the dispatched-eval slot, the baseline cache, and the pre-budget lookup that reduces the charge to one main eval.
  • Cross-run baseline cache entries written without these fields are misses (re-measured once), never hits.

Compatibility (RELEASING.md; measurement caches):

  • Surfaces: the dispatched-eval slot and scheduler names (via the determinant), and baseline cache entries.
  • In-flight runs: a run parked with an eval written or running under the old slot adopts it. When the new slot is absent, the measurer looks for the legacy slot in the same run directory and uses it, so nothing is re-dispatched and nothing is charged twice. A current slot always takes precedence. A run's GPU count is fixed within the run, so this does not reopen the cross-run bug.
  • Fixtures: tests/fixtures/dispatched_pre_gpu_identity/, written by the previous kernel, with a completed result and an in-flight job. Each is exercised through the new code: first pass, an idempotent second pass, a retry after an interruption, and a missing submit marker.
  • Rollback: safe. The old kernel ignores the extra baseline fields and finds its own slots for evals written before the upgrade. Evals written by the new kernel are re-dispatched by an old kernel.
  • Upgrading: line in the CHANGELOG.

Built by Codex. I reviewed it and asked for the in-flight adoption and fixtures. Gate: pytest, ruff check, ruff format --check, mypy.

…r slot

The dispatched-eval determinant omitted the GPU count, so a result could
be reused across GPU counts and discount a gate's charge. The determinant,
the baseline cache and the pre-budget lookup now carry the GPU count and
the lane's GPU type. A run adopts an eval written under the previous slot
name in its own directory, so nothing in flight is re-dispatched.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 9d8df6a1 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: 1 blocking, 0 advisory.

1 finding attached to the lines below.

Merged verdict: blocking legacy slot adoption can reuse an in-flight evaluation across a GPU lane type change. No findings rejected; the credentials, deployment, and lifecycle reports are one duplicate finding consolidated at the sharpest location and severity.

Comment thread src/outerloop/measure.py Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head e25c1194 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye
renmengye merged commit 9716372 into main Oct 2, 2026
5 checks passed
@renmengye
renmengye deleted the fix/eval-cache-gpu-count branch October 2, 2026 18:30
@renmengye renmengye mentioned this pull request Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant