feat(bench): add an mcp tool level beside bin for the agent benchmark - #492
Merged
Merged
Conversation
Teakowa
added this pull request to stack #495
October 3, 2026 17:33
e54-bot
force-pushed
the
feat/474-mcp-level
branch
8 times, most recently
from
October 4, 2026 06:02
1367c58 to
270200e
Compare
Adds `level` to the condition grid so the same scenario, prompt, grader, model, and trial can run with Wright as CLI (bin, the default) or as native MCP tools (mcp). Under mcp the wright CLI stays off PATH and the harness hands the adapter BENCH_MCP_CMD — `wright serve --transport mcp` through the existing tracing shim — while the canary enforces that the CLI is unreachable. - agent_bench: LEVELS + --level flag, cell normalization/labels (wright-mcp+.../knowledge/network), wright-only level validation, BENCH_TOOL_LEVEL/BENCH_MCP_CMD env, mcp no-server invalid check, and result.toolCalls from the normalized transcript. - bench_trace: serve lines now carry their session id and transport; serve_request mirrors each transport's answer rule (blank lines and JSON-RPC notifications are silent, everything else is answered), so request/response pairing cannot desynchronize and concurrent sessions do not cross-pair. tools/call maps to the wright op its tool name carries; handshake methods keep mcp: names — excluded from uses but their response bytes (the tools/list schema payload) count as output. Structured refusals are refused, only parse-level errors are malformed. - direct adapter: stdio JSON-RPC client that initializes, lists tools, converts schemas to model tools, dispatches calls, and closes reliably; the module state resets per run and setup failures exit 1 (agent error), not provider-interrupted. - bench_report: paired mcp-vs-bin section on (scenario, agent, trial, knowledge, network, skills) with Wilson intervals on usable/passed, search/read counts, wright/bash/tool calls, turns, and tokens. Closes #474
Mcp._lines polled select() on the pipe FD but read with TextIOWrapper.readline(), which buffers ahead past what select reports: a line behind the matching reply sat in the wrapper's buffer while the next request waited on an empty FD and timed out. Raw os.read into self.pending keeps every complete line either yielded or pending across requests; a line at the buffer boundary and a final line without a newline are handled the same way.
… before forwarding serve_request reused DECISION_COMMANDS for the jsonrpc method map, which has lint and lacks compile: a real compile call classified as transport traffic (invisible to toolUse and friction), while a lint the server can only answer MethodNotFound counted as a Wright use. serve.rs accepts compile/check/analyze/inspect, so those get their own constant. serve_tee also logged the req event after writing to the server, letting a fast response reach the trace first and desynchronizing serve_pairs; it now logs before the write.
Teakowa
force-pushed
the
feat/474-mcp-level
branch
from
October 4, 2026 06:06
270200e to
dbf02e1
Compare
Teakowa
approved these changes
Oct 4, 2026
bench_report.load() also parsed result.json unguarded, so a truncated file left by a killed run still crashed evaluate's report step for trials the new run never attempted. The atomic writer moves to bench_report — the one module every artifact reader already imports — and now also covers score.json, leaderboard.json, and summary.json. compare() and load_entries() treat a partial score.json like a run whose scoring never finished instead of crashing on JSONDecodeError.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes #474. Adds
levelto the benchmark condition grid so the same scenario, prompt, grader, model, and trial can run with Wright as a traced CLI (bin, the default — unchanged) or as native MCP tools (mcp). Stacked on #491.LEVELS = ("bin", "mcp"); labelwright-mcp[+skills]/knowledge/network;levelflows throughrun --level, matrix cell dicts, normalization, validation (mcpis awright-only level), and result metadata (condition.level,BENCH_TOOL_LEVEL).mcpthewrightCLI stays offPATH; the adapter getsBENCH_MCP_CMD—wright serve --transport mcp <workspace>wrapped in the same tracing shim — and the canary fails the run ifwrightis reachable. A run where the adapter never started the server isinvalid(skipped when the attempt was alreadyprovider-interrupted; invalid reasons now accumulate instead of overwriting).initialize,notifications/initialized,tools/list→ native tool schemas →tools/calldispatch → results back to the model. Per-request read bound (COMMAND_SECONDS), reliable close, module-global client reset eachmain()so in-process runs can't leak a stale server. Setup failures exit 1 (agent error), notprovider-interrupted.serve_requestmirrors each transport's real answer rule — blank lines and JSON-RPC notifications are answered with silence, everything else (including unparseable or invalid input) gets an error response. Pairing is FIFO per session, so silent lines can't desynchronize requests/responses and concurrent sessions can't cross-pair.tools/callmapswright_*tool names back to Wright ops;mcp:*handshake methods are excluded from uses but their response bytes (thetools/listschema payload) count inoutputTokensEstimate. Structured rejections classify asrefused; only-32700/-32600/malformed-requestaremalformed.Level comparisonpairsmcp/binon (scenario, agent, trial, knowledge, network, skills) with Wilson intervals on usable and passed rates plussearchReads(wright invocations + search/readbashcalls), wright/bash/tool calls, turns, tokens/run, and paired gain/loss + token saving.Test plan
python3 -m unittest discoverinbenchmarks/agent— 105 tests, all pass (WRIGHT_BINset so shim tests run)mcp;serve_requestanswer rule per transport; desync and concurrent-session pairing;serve_errorclasses;call_counts/shell_search_reads; adapter MCP register/dispatch/close/failure paths; report pairing + metricsagent_bench.py validatecleanrepair-runaway-loopPASS at both levels; mcpwrightUse= 4 invocations (project,symbols,check×2), handshake separated,mcp:tools/list≈ 1284 est. output tokens,failedInvocations0, E11 pass; paired report renders with Wilson intervalsgit diff --checkcleanserve.rs/mcp.rsanswer rules; missing unit coverage), minors fixed (accumulated invalid reasons, INFRA gating,cell.get("level")consistency, request timeout, non-dict transcript input,mkdtempsetup, docs/example drift); unreachableMCP_ADAPTERSpreflight gate removed rather than left deadKnown limits: only the
directadapter registersBENCH_MCP_CMDtoday (the docs name it; other adapters' mcp runs come backinvalid, which is the honest signal).