feat(bench): standardized Wright Agent Score for model comparison - #491
Merged
Merged
Conversation
6 tasks done
Teakowa
added this pull request to stack #495
October 3, 2026 17:33
…p actually observed
…ck off retries, and grade a missing entry as an agent failure
…flight, and a dry run
… sibling runs and harness data
…l id, run sequentially by default, and write refreshed logins back
Independent review of the v3 score branch found a verbatim duplicate of the credential block, per-adapter copies of the transient-failure classification, crashes when --effort is omitted, preflight that missed unusable sandboxes and missing logins, a KeyError on results without a status, and a leaderboard that could rank a one-track run above full two-track runs. - share INFRA_EXIT and TRANSIENT through adapters/common.py - tolerate a missing BENCH_THINKING in the agy and codex adapters - preflight the file sandbox and the adapter's primary login - classify any non-completed, non-timeout status as excluded - rank only runs covering every language track; partial runs are listed as not comparable - derive leaderboard prose and network claims from the actual runs - exclude BUILD.json from its own skill hash so verification is stable - fail the suite exit code when a model finishes with errors
The second review found that TRANSIENT was matched against the agent's own output in claude_code, devin, and grok: a failed run whose message mentioned "timeout" or "quota" would exit 75, be recorded as provider-interrupted, and be excluded from the score, inflating it. Only provider channels (stderr, structured error fields) now classify. - devin.transient takes stdout and stderr separately; the anchored model-catalog error still counts on either stream - grok and pi classify on the structured error field only - claude_code checks stderr only and names its plugin dir predictably - every file an adapter unconditionally reads is in CREDENTIALS, so preflight names it; pi's optional antigravity login moves to OPTIONAL_CREDENTIALS - stream parsers skip non-dict JSON and use .get() throughout; codex throttles its per-line session rescan - matrix config options can no longer inject a silent deny_read - result loading drops records missing the fields renderers need; score cards sort mixed recorded/unrecorded enforcement values - the HTML page lists non-comparable runs like the Markdown does
Follow-up to the previous fix, from the same review pass: nested event fields could still be truthy non-dicts (a str is truthy but has no .get), one direct index survived (agy step_index), and a retried devin attempt could read the previous attempt's export file. - add as_dict() in adapters/common and use it for every nested event container the stream loops dereference - clear devin-export.json between infrastructure retries - tighten the devin catalog regex to one line and check each stream's tail independently
…lure and read the effort from the model id
… reproduces the run The manifest only carried skill_dirs and wiki_dir, so 'agent_bench.py matrix <run>/matrix.json' dropped every trial-time setting the run was made with: it ran unsandboxed, wrote into --out instead of the run directory, and lost env_pass (HOME or the direct API keys), credentials, allow/deny lists, timeout, canary, ancestor check, retry policy, adapter, and wright. The options dict now records each option at its effective value; out and out_root serialize as "." and "..", which the loader resolves against the matrix file's directory, so the file stays a self-contained, relocatable manifest of its run. trial_dir honors options.out, and skill_dirs/wiki_dir/out/out_root resolve relative to the file.
…/network enforcement identity_of covered the binary, skills, suite, agent, and protocol but not how the agent's filesystem and network were policed. A directory holding both sandboxed and unrestricted trials — which an unsandboxed 'matrix' rerun used to produce — still emitted a card, silently mixing trials that could read the answer keys with ones that could not. The card now records and discloses the read policy (allow-list mode, not the per-trial path lists) and compare() warns when runs differ.
…e claude-code sandbox login note sync_credentials_back wrote the staging file at the default umask before chmod; it is now created at 0o600 (replacing any leftover first). evaluate prints a one-line note that claude-code can read but not refresh its login under the file sandbox, and the vestigial single-element argument loop over the suite parser is flattened.
…trix
Relative --skill-dir/--wiki-dir/--allow-read/--deny-read values were written to matrix.json raw, then resolved against the run directory on replay instead of the evaluate working directory, so a reproduced run looked for skills, wiki, and read policies under the wrong root. Serialize them resolved to absolute paths; the generated out/out_root stay manifest-relative ('.'/'..') so the run directory can still move or be archived.
On the read side the same manifest-relative rule now covers allow_read/deny_read entries and ~ expansion, and a null path option means 'unset' rather than overriding with a value that crashes path handling. docs/agent-benchmark.md now names every path-valued option, documents fileReadEnforcement/fileWriteEnforcement, and states that mixed-enforcement runs refuse a score card.
…rded' compare() indexed every COMPARABLE key on each card's identity, so a score.json written before file-read/write/network enforcement joined the identity crashed with KeyError instead of warning that the cards differ. ident.get(k) reports the absent field as a difference, which is the same treatment card() gives unrecorded enforcement.
Review nits on the reproduction fix: --skill-dir/--wiki-dir now expand ~ like the read policies do; wright follows the same manifest-relative rule as the other path options so a hand-authored relative binary resolves against the file instead of the trial workspace; a null skill_dirs means unset like the other path keys instead of crashing on .items(); compare() reads agent/effort with .get for uniformly missing-field-tolerant old cards; and evaluate rejects --name values that are not a single directory component ('..', 'a/b', empty), which previously let the run directory escape --out. The reproduction test also covers relative --wiki-dir and infra_retries overrides.
…stdin direct.py parsed API replies with raw indexing, so a malformed payload raised KeyError/IndexError/JSONDecodeError as a traceback instead of a classified failure; both providers now wrap response-shape reads as a non-transient ProviderError. The opencode, pi, codex, and claude-code adapters also wrote the whole prompt to stdin before draining the pipes, which deadlocks on a prompt larger than the pipe buffer against a chatty CLI — they now share common.feed_stdin, which feeds stdin on a thread.
…box profile everywhere read_policy allowed benchmarks/agent/oracle unconditionally, so a wright or none cell could read and run the pinned upstream compiler — a grading authority — blurring the tool-condition contrast the sandbox exists to enforce. It is readable only when BENCH_TOOL is overpy, whose launcher execs it. The generated agent.sb profile sat inside the readable run directory and listed the hidden paths; the profile now ends with a deny rule for itself.
cmd_suite filtered with startswith, so --only codex:gpt-6 also selected codex:gpt-6-luna. Selection is now exact against the adapter name or the adapter:model pair.
the adapters that piped stderr read it only after stdout closed, so a child that wrote more than a pipe buffer of stderr (or any stderr before consuming a large stdin prompt) blocked forever. drain() reads the pipe on a thread and returns its text after wait(), the same shape as feed_stdin; codex/agy already redirect stderr to a log file and devin/direct use run()/communicate().
…r another name the read deny on agent.sb was literal-path only, and the profile sits inside the writable run dir, so the agent could rename or hardlink it to a name the deny does not cover and read the hidden-path list. the profile now also denies file-write on its own literal (last matching rule wins over the run-dir allow, blocking rename, unlink, overwrite, and the chflags that would clear the lock), and the harness sets UF_IMMUTABLE on it for the trial, which makes hardlinking fail outright where path rules cannot reach. the sandbox probes now try the rename and link escapes themselves.
a run killed between setting agent.sb's immutable flag and clearing it left rmtree unable to remove the run dir, and the next trial failed at mkdir. the agent itself can also chflags or chmod anything inside its writable tree. drop() now restores writability at the failing node and retries; rmtree's onexc does the same inside each tree.
retrying the failed op in place broke on traversal callbacks like os.open, which take more than the path; and chflags/chmod follow symlinks by default, so a planted stale symlink could steer the repairs onto a host file. leaf removals still retry immediately, a blocked traversal is repaired and picked up by the next pass, and the repairs no longer follow links.
- wright_mismatch returns early when the binary is missing, so a nonexistent --wright reports through preflight's "cannot start" list instead of crashing evaluate/suite with FileNotFoundError - result.json is written atomically (temp + os.replace); a partial file left by a mid-write kill is treated as unfinished and retried by matrix instead of crashing pool.map with JSONDecodeError — the same tolerance now covers wright_mismatch's environment scan - --suite-name takes the same single-directory-name check as --name, closing the ".." escape under --out - file_sha256's signature admits the Path callers actually pass
Teakowa
approved these changes
Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements the
wright-agent-bench/v3contract: a standardized, versioned Wright Agent Score so coding agents can be compared on real Workshop and OverPy tasks under one canonical condition.usableoutcomes across scenarios; 95% CI from a two-stage percentile bootstrap (resample scenarios, then trials within them), 10,000 draws, seed 467.pass^kstays a secondary metric.directadapter behind the sameBENCH_*contract, so agent identity (program + model + effort) is what's scored.score.json/RESULTS.md,compareacross runs, and aleaderboardpage (Markdown/HTML/JSON) that ranks only environment-identical, full-coverage runs and lists the rest under "Not comparable".Verification
python3 -m unittest discover -p 'test_*.py'inbenchmarks/agentwithWRIGHT_BINset: 92 tests, OK.swe-2-max, 48/48 trials, 0 exclusions: Workshop 62.5 (95% CI 29.2–91.7), OverPy 62.5 (25.0–87.5), Pass^3 50.0 / 62.5.git diff --checkandcargo fmt --all -- --checkclean; the branch touches no Rust code.--effort, missing sandbox/credential preflight, and partial-track ranking; the second found and we fixed a score-critical misclassification where agent-authored text containing "timeout"/"quota" could mark a failed runprovider-interruptedand be excluded — only provider channels classify now.Disclosed limits
declared-only— agents were instructed not to use the network, not blocked; the score card and leaderboard say so explicitly rather than claiming the canonical condition.gpt-6-luna) hit the provider's usage limit at ~30/48 trials. The harness classified the interruption asprovider-interrupted, excluded it from outcomes, and the run resumes from where it stopped when quota returns.Closes #467