refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303
Merged
Merged
Conversation
Astro-Han
force-pushed
the
refactor/benchmark-identity-out-of-dist
branch
from
August 6, 2026 02:12
1a563c6 to
2a08231
Compare
Astro-Han
changed the base branch from
main
to
fix/harness-maka-repo-mount-scope
August 6, 2026 02:12
Astro-Han
force-pushed
the
refactor/benchmark-identity-out-of-dist
branch
2 times, most recently
from
August 6, 2026 06:04
e79c5ba to
683e1c9
Compare
Astro-Han
changed the base branch from
fix/harness-maka-repo-mount-scope
to
main
August 6, 2026 06:04
…aded arm mounts An arm executes this package's dist, so a benchmark's revision, upstream URL and task list compiled into a shipped module are readable from inside the container being scored — which is how the #1970 run turned a task into a lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval and src; the identity itself still shipped in harness-ab-manifest.js, and the path-level mount checks cannot see it because they inspect declared paths, never what the mounted files contain. The data moves to harbor/benchmark-identity.json, which nothing mounts, and the assertions in src become the comparison without the answer. A content test over the shipped dist now holds the invariant directly; it caught a Terminal-Bench task id hardcoded as a smoke default that the manifest was already supplying. Collapses two identical task-source branches: with each profile carrying its own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held. The content check read only the top level of dist while the whole directory is mounted, and described the compiled tests as excluded when they are not. It now scans dist recursively and states two boundaries, because they are two different things: the retrieval keys — revision, upstream URL, tree fingerprint — must be absent from every mounted file, while a single task id is not a retrieval key and is only held out of the shipped modules, where a complete list could accumulate. No mount list may declare the identity file, the Maka one included, and the DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot The scan filtered to `.js`, so a source map — whose `sourcesContent` carries the whole of `src` — would have walked straight past it. And the retrieval keys named every task-tree fingerprint but one.
Astro-Han
force-pushed
the
refactor/benchmark-identity-out-of-dist
branch
from
August 6, 2026 06:11
683e1c9 to
d0e7cf3
Compare
Astro-Han
marked this pull request as ready for review
August 6, 2026 06:25
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
A graded arm executes
packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored.harness-ab-manifest.tscarried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.
#2298 narrowed the mount so
docs/evalandsrcno longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.How
packages/headless/harbor/benchmark-identity.json, loaded byharbor/benchmark-identity.mjs. Neither is mounted by any arm —agent-repo-mount.tsnames exactly one file out ofharbor/, and it isrun-cell.mjs.harness-ab-manifest.tskeeps the assertions and loses the answers:assertTerminalBench21TaskSetand its three siblings becomeassertFrozenTaskSet(label, expected, actual)andassertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.HARNESS_BENCHMARK_PROFILESnow carries its owntaskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.harbor/*.mjsentrypoints and by tests.Verification
A new test in
harness-ab-manifest.test.tsreads every shippeddist/*.jsand asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak:harbor-smoke-config.tshardcoded'*sqlite-with-gcov'— a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.Red/green confirmed: restoring that literal fails the new test.
Loading the real entrypoint reproduces all three profiles byte-for-byte:
npm run -w @maka/headless test— 1422 pass, 0 fail. Lint and format clean.Known residue
Compiled tests under
dist/__tests__are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out ofdist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.dist/terminal-bench-adapter.jsremains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.