Skip to content

feat(bench): add an mcp tool level beside bin for the agent benchmark - #492

Merged
Teakowa merged 4 commits into
mainfrom
feat/474-mcp-level
Oct 4, 2026
Merged

Teakowa merged 4 commits into
mainfrom
feat/474-mcp-level

Conversation

@e54-bot

@e54-bot e54-bot commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

Closes #474. Adds level to the benchmark condition grid so the same scenario, prompt, grader, model, and trial can run with Wright as a traced CLI (bin, the default — unchanged) or as native MCP tools (mcp). Stacked on #491.

  • Cell model: LEVELS = ("bin", "mcp"); label wright-mcp[+skills]/knowledge/network; level flows through run --level, matrix cell dicts, normalization, validation (mcp is a wright-only level), and result metadata (condition.level, BENCH_TOOL_LEVEL).
  • Environment: under mcp the wright CLI stays off PATH; the adapter gets BENCH_MCP_CMD — wright serve --transport mcp <workspace> wrapped in the same tracing shim — and the canary fails the run if wright is reachable. A run where the adapter never started the server is invalid (skipped when the attempt was already provider-interrupted; invalid reasons now accumulate instead of overwriting).
  • direct adapter: stdio JSON-RPC client — initialize, notifications/initialized, tools/list → native tool schemas → tools/call dispatch → results back to the model. Per-request read bound (COMMAND_SECONDS), reliable close, module-global client reset each main() so in-process runs can't leak a stale server. Setup failures exit 1 (agent error), not provider-interrupted.
  • Tracing: serve lines now carry a session id and their transport, and serve_request mirrors each transport's real answer rule — blank lines and JSON-RPC notifications are answered with silence, everything else (including unparseable or invalid input) gets an error response. Pairing is FIFO per session, so silent lines can't desynchronize requests/responses and concurrent sessions can't cross-pair. tools/call maps wright_* tool names back to Wright ops; mcp:* handshake methods are excluded from uses but their response bytes (the tools/list schema payload) count in outputTokensEstimate. Structured rejections classify as refused; only -32700/-32600/malformed-request are malformed.
  • Report: Level comparison pairs mcp/bin on (scenario, agent, trial, knowledge, network, skills) with Wilson intervals on usable and passed rates plus searchReads (wright invocations + search/read bash calls), wright/bash/tool calls, turns, tokens/run, and paired gain/loss + token saving.

Test plan

  • python3 -m unittest discover in benchmarks/agent — 105 tests, all pass (WRIGHT_BIN set so shim tests run)
  • New unit coverage: cell/label/env/canary for mcp; serve_request answer rule per transport; desync and concurrent-session pairing; serve_error classes; call_counts/shell_search_reads; adapter MCP register/dispatch/close/failure paths; report pairing + metrics
  • agent_bench.py validate clean
  • Real end-to-end evidence (scripted agent, real harness + real wright binary): repair-runaway-loop PASS at both levels; mcp wrightUse = 4 invocations (project, symbols, check×2), handshake separated, mcp:tools/list ≈ 1284 est. output tokens, failedInvocations 0, E11 pass; paired report renders with Wilson intervals
  • git diff --check clean
  • Independent review pass: majors fixed (serve-pairing desync on silent lines — verified against serve.rs/mcp.rs answer rules; missing unit coverage), minors fixed (accumulated invalid reasons, INFRA gating, cell.get("level") consistency, request timeout, non-dict transcript input, mkdtemp setup, docs/example drift); unreachable MCP_ADAPTERS preflight gate removed rather than left dead

Known limits: only the direct adapter registers BENCH_MCP_CMD today (the docs name it; other adapters' mcp runs come back invalid, which is the honest signal).

@Teakowa
Teakowa added this pull request to stack #495 October 3, 2026 17:33
@e54-bot
e54-bot force-pushed the feat/474-mcp-level branch 8 times, most recently from 1367c58 to 270200e Compare October 4, 2026 06:02
Base automatically changed from feat/467-v3 to main October 4, 2026 06:06
Adds `level` to the condition grid so the same scenario, prompt, grader,
model, and trial can run with Wright as CLI (bin, the default) or as
native MCP tools (mcp). Under mcp the wright CLI stays off PATH and the
harness hands the adapter BENCH_MCP_CMD — `wright serve --transport mcp`
through the existing tracing shim — while the canary enforces that the
CLI is unreachable.

- agent_bench: LEVELS + --level flag, cell normalization/labels
  (wright-mcp+.../knowledge/network), wright-only level validation,
  BENCH_TOOL_LEVEL/BENCH_MCP_CMD env, mcp no-server invalid check, and
  result.toolCalls from the normalized transcript.
- bench_trace: serve lines now carry their session id and transport;
  serve_request mirrors each transport's answer rule (blank lines and
  JSON-RPC notifications are silent, everything else is answered), so
  request/response pairing cannot desynchronize and concurrent sessions
  do not cross-pair. tools/call maps to the wright op its tool name
  carries; handshake methods keep mcp: names — excluded from uses but
  their response bytes (the tools/list schema payload) count as output.
  Structured refusals are refused, only parse-level errors are malformed.
- direct adapter: stdio JSON-RPC client that initializes, lists tools,
  converts schemas to model tools, dispatches calls, and closes reliably;
  the module state resets per run and setup failures exit 1 (agent error),
  not provider-interrupted.
- bench_report: paired mcp-vs-bin section on (scenario, agent, trial,
  knowledge, network, skills) with Wilson intervals on usable/passed,
  search/read counts, wright/bash/tool calls, turns, and tokens.

Closes #474
Mcp._lines polled select() on the pipe FD but read with TextIOWrapper.readline(), which buffers ahead past what select reports: a line behind the matching reply sat in the wrapper's buffer while the next request waited on an empty FD and timed out. Raw os.read into self.pending keeps every complete line either yielded or pending across requests; a line at the buffer boundary and a final line without a newline are handled the same way.
… before forwarding

serve_request reused DECISION_COMMANDS for the jsonrpc method map, which has lint and lacks compile: a real compile call classified as transport traffic (invisible to toolUse and friction), while a lint the server can only answer MethodNotFound counted as a Wright use. serve.rs accepts compile/check/analyze/inspect, so those get their own constant. serve_tee also logged the req event after writing to the server, letting a fast response reach the trace first and desynchronizing serve_pairs; it now logs before the write.
@Teakowa
Teakowa force-pushed the feat/474-mcp-level branch from 270200e to dbf02e1 Compare October 4, 2026 06:06
bench_report.load() also parsed result.json unguarded, so a truncated
file left by a killed run still crashed evaluate's report step for
trials the new run never attempted. The atomic writer moves to
bench_report — the one module every artifact reader already imports —
and now also covers score.json, leaderboard.json, and summary.json.
compare() and load_entries() treat a partial score.json like a run
whose scoring never finished instead of crashing on JSONDecodeError.
@Teakowa
Teakowa merged commit d3e4f1d into main Oct 4, 2026
15 checks passed
@Teakowa
Teakowa deleted the feat/474-mcp-level branch October 4, 2026 07:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

Add an mcp level to the agent benchmark

2 participants