Skip to content

feat(bench): standardized Wright Agent Score for model comparison - #491

Merged
Teakowa merged 32 commits into
mainfrom
feat/467-v3
Oct 4, 2026
Merged

Teakowa merged 32 commits into
mainfrom
feat/467-v3

Conversation

@e54-bot

@e54-bot e54-bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements the wright-agent-bench/v3 contract: a standardized, versioned Wright Agent Score so coding agents can be compared on real Workshop and OverPy tasks under one canonical condition.

  • Two headline tracks — Wright Workshop Agent Score and Wright OPY Agent Score, never pooled into one number.
  • Canonical condition — Wright tool + Wright skill, no extra knowledge, network off, held-out test split (8 scenarios per language, 3 trials each).
  • Score — macro-average of deterministic binary usable outcomes across scenarios; 95% CI from a two-stage percentile bootstrap (resample scenarios, then trials within them), 10,000 draws, seed 467. pass^k stays a secondary metric.
  • Reproducible identity — every result records the Wright binary sha256, skill content hashes, suite hash, adapter/model/effort, and harness commit; scoring and comparison refuse or flag runs made under different environments.
  • Adapters — devin, codex, pi, agy, opencode, grok, claude-code, and a harness-free direct adapter behind the same BENCH_* contract, so agent identity (program + model + effort) is what's scored.
  • Isolation and reliability — macOS file sandbox (trial-directory-only writes, answer keys / sibling runs / harness data hidden), per-adapter credential copies with write-back of refreshed logins, provider-failure classification (exit 75) with backoff and full resume.
  • Publishable output — per-run score.json/RESULTS.md, compare across runs, and a leaderboard page (Markdown/HTML/JSON) that ranks only environment-identical, full-coverage runs and lists the rest under "Not comparable".

Verification

  • python3 -m unittest discover -p 'test_*.py' in benchmarks/agent with WRIGHT_BIN set: 92 tests, OK.
  • Real end-to-end run — Devin swe-2-max, 48/48 trials, 0 exclusions: Workshop 62.5 (95% CI 29.2–91.7), OverPy 62.5 (25.0–87.5), Pass^3 50.0 / 62.5.
  • git diff --check and cargo fmt --all -- --check clean; the branch touches no Rust code.
  • Two independent review passes, both clean after fixes: the first found the duplicated credential block, adapter crashes without --effort, missing sandbox/credential preflight, and partial-track ranking; the second found and we fixed a score-critical misclassification where agent-authored text containing "timeout"/"quota" could mark a failed run provider-interrupted and be excluded — only provider channels classify now.

Disclosed limits

  • The evidence run's network enforcement is declared-only — agents were instructed not to use the network, not blocked; the score card and leaderboard say so explicitly rather than claiming the canonical condition.
  • A second-model comparison run (Pi gpt-6-luna) hit the provider's usage limit at ~30/48 trials. The harness classified the interruption as provider-interrupted, excluded it from outcomes, and the run resumes from where it stopped when quota returns.

Closes #467

@Teakowa
Teakowa added this pull request to stack #495 October 3, 2026 17:33
Teakowa added 28 commits October 4, 2026 13:52
…ck off retries, and grade a missing entry as an agent failure
…l id, run sequentially by default, and write refreshed logins back
Independent review of the v3 score branch found a verbatim duplicate of
the credential block, per-adapter copies of the transient-failure
classification, crashes when --effort is omitted, preflight that missed
unusable sandboxes and missing logins, a KeyError on results without a
status, and a leaderboard that could rank a one-track run above full
two-track runs.

- share INFRA_EXIT and TRANSIENT through adapters/common.py
- tolerate a missing BENCH_THINKING in the agy and codex adapters
- preflight the file sandbox and the adapter's primary login
- classify any non-completed, non-timeout status as excluded
- rank only runs covering every language track; partial runs are listed
  as not comparable
- derive leaderboard prose and network claims from the actual runs
- exclude BUILD.json from its own skill hash so verification is stable
- fail the suite exit code when a model finishes with errors
The second review found that TRANSIENT was matched against the agent's
own output in claude_code, devin, and grok: a failed run whose message
mentioned "timeout" or "quota" would exit 75, be recorded as
provider-interrupted, and be excluded from the score, inflating it.
Only provider channels (stderr, structured error fields) now classify.

- devin.transient takes stdout and stderr separately; the anchored
  model-catalog error still counts on either stream
- grok and pi classify on the structured error field only
- claude_code checks stderr only and names its plugin dir predictably
- every file an adapter unconditionally reads is in CREDENTIALS, so
  preflight names it; pi's optional antigravity login moves to
  OPTIONAL_CREDENTIALS
- stream parsers skip non-dict JSON and use .get() throughout;
  codex throttles its per-line session rescan
- matrix config options can no longer inject a silent deny_read
- result loading drops records missing the fields renderers need;
  score cards sort mixed recorded/unrecorded enforcement values
- the HTML page lists non-comparable runs like the Markdown does
Follow-up to the previous fix, from the same review pass: nested event
fields could still be truthy non-dicts (a str is truthy but has no
.get), one direct index survived (agy step_index), and a retried devin
attempt could read the previous attempt's export file.

- add as_dict() in adapters/common and use it for every nested event
  container the stream loops dereference
- clear devin-export.json between infrastructure retries
- tighten the devin catalog regex to one line and check each stream's
  tail independently
… reproduces the run

The manifest only carried skill_dirs and wiki_dir, so 'agent_bench.py matrix <run>/matrix.json' dropped every trial-time setting the run was made with: it ran unsandboxed, wrote into --out instead of the run directory, and lost env_pass (HOME or the direct API keys), credentials, allow/deny lists, timeout, canary, ancestor check, retry policy, adapter, and wright.

The options dict now records each option at its effective value; out and out_root serialize as "." and "..", which the loader resolves against the matrix file's directory, so the file stays a self-contained, relocatable manifest of its run. trial_dir honors options.out, and skill_dirs/wiki_dir/out/out_root resolve relative to the file.
…/network enforcement

identity_of covered the binary, skills, suite, agent, and protocol but not how the agent's filesystem and network were policed. A directory holding both sandboxed and unrestricted trials — which an unsandboxed 'matrix' rerun used to produce — still emitted a card, silently mixing trials that could read the answer keys with ones that could not. The card now records and discloses the read policy (allow-list mode, not the per-trial path lists) and compare() warns when runs differ.
…e claude-code sandbox login note

sync_credentials_back wrote the staging file at the default umask before chmod; it is now created at 0o600 (replacing any leftover first). evaluate prints a one-line note that claude-code can read but not refresh its login under the file sandbox, and the vestigial single-element argument loop over the suite parser is flattened.
…trix

Relative --skill-dir/--wiki-dir/--allow-read/--deny-read values were written to matrix.json raw, then resolved against the run directory on replay instead of the evaluate working directory, so a reproduced run looked for skills, wiki, and read policies under the wrong root. Serialize them resolved to absolute paths; the generated out/out_root stay manifest-relative ('.'/'..') so the run directory can still move or be archived.

On the read side the same manifest-relative rule now covers allow_read/deny_read entries and ~ expansion, and a null path option means 'unset' rather than overriding with a value that crashes path handling. docs/agent-benchmark.md now names every path-valued option, documents fileReadEnforcement/fileWriteEnforcement, and states that mixed-enforcement runs refuse a score card.
…rded'

compare() indexed every COMPARABLE key on each card's identity, so a score.json written before file-read/write/network enforcement joined the identity crashed with KeyError instead of warning that the cards differ. ident.get(k) reports the absent field as a difference, which is the same treatment card() gives unrecorded enforcement.
Review nits on the reproduction fix: --skill-dir/--wiki-dir now expand ~ like the read policies do; wright follows the same manifest-relative rule as the other path options so a hand-authored relative binary resolves against the file instead of the trial workspace; a null skill_dirs means unset like the other path keys instead of crashing on .items(); compare() reads agent/effort with .get for uniformly missing-field-tolerant old cards; and evaluate rejects --name values that are not a single directory component ('..', 'a/b', empty), which previously let the run directory escape --out. The reproduction test also covers relative --wiki-dir and infra_retries overrides.
…stdin

direct.py parsed API replies with raw indexing, so a malformed payload raised KeyError/IndexError/JSONDecodeError as a traceback instead of a classified failure; both providers now wrap response-shape reads as a non-transient ProviderError. The opencode, pi, codex, and claude-code adapters also wrote the whole prompt to stdin before draining the pipes, which deadlocks on a prompt larger than the pipe buffer against a chatty CLI — they now share common.feed_stdin, which feeds stdin on a thread.
…box profile everywhere

read_policy allowed benchmarks/agent/oracle unconditionally, so a wright or none cell could read and run the pinned upstream compiler — a grading authority — blurring the tool-condition contrast the sandbox exists to enforce. It is readable only when BENCH_TOOL is overpy, whose launcher execs it. The generated agent.sb profile sat inside the readable run directory and listed the hidden paths; the profile now ends with a deny rule for itself.
cmd_suite filtered with startswith, so --only codex:gpt-6 also selected codex:gpt-6-luna. Selection is now exact against the adapter name or the adapter:model pair.
the adapters that piped stderr read it only after stdout closed, so a child that wrote more than a pipe buffer of stderr (or any stderr before consuming a large stdin prompt) blocked forever. drain() reads the pipe on a thread and returns its text after wait(), the same shape as feed_stdin; codex/agy already redirect stderr to a log file and devin/direct use run()/communicate().
…r another name

the read deny on agent.sb was literal-path only, and the profile sits inside the writable run dir, so the agent could rename or hardlink it to a name the deny does not cover and read the hidden-path list. the profile now also denies file-write on its own literal (last matching rule wins over the run-dir allow, blocking rename, unlink, overwrite, and the chflags that would clear the lock), and the harness sets UF_IMMUTABLE on it for the trial, which makes hardlinking fail outright where path rules cannot reach. the sandbox probes now try the rename and link escapes themselves.
a run killed between setting agent.sb's immutable flag and clearing it left rmtree unable to remove the run dir, and the next trial failed at mkdir. the agent itself can also chflags or chmod anything inside its writable tree. drop() now restores writability at the failing node and retries; rmtree's onexc does the same inside each tree.
retrying the failed op in place broke on traversal callbacks like os.open, which take more than the path; and chflags/chmod follow symlinks by default, so a planted stale symlink could steer the repairs onto a host file. leaf removals still retry immediately, a blocked traversal is repaired and picked up by the next pass, and the repairs no longer follow links.
- wright_mismatch returns early when the binary is missing, so a
  nonexistent --wright reports through preflight's "cannot start" list
  instead of crashing evaluate/suite with FileNotFoundError
- result.json is written atomically (temp + os.replace); a partial file
  left by a mid-write kill is treated as unfinished and retried by
  matrix instead of crashing pool.map with JSONDecodeError — the same
  tolerance now covers wright_mismatch's environment scan
- --suite-name takes the same single-directory-name check as --name,
  closing the ".." escape under --out
- file_sha256's signature admits the Path callers actually pass
@Teakowa
Teakowa merged commit 1ad68ec into main Oct 4, 2026
3 checks passed
@Teakowa
Teakowa deleted the feat/467-v3 branch October 4, 2026 06:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

Define a standardized Wright Agent Score for model comparison

2 participants