Problem
There is no measurement of whether native Wright tools improve on shell + JSON for task success and search/read/tool-call efficiency. See ADR-0020 (Proposed in the ADR PR for #470).
Scope
Add mcp as a wright level beside bin in the agent benchmark condition grid (<wright>/<knowledge>/<network>), with scenario, prompt, grader, model, and trial identity unchanged:
- the harness describes the level to the adapter, and the reference adapter registers the
wright serve --transport mcp server for the agent;
- the Wright trace records MCP tool calls so
wrightUse, friction, and the report cover them;
- the report compares
mcp with bin on usable/passed rate with Wilson intervals and on search/read/tool-call counts, turns, and tokens; tool-schema context counts toward the mcp cost;
docs/agent-benchmark.md documents the level.
Non-goals
Acceptance criteria
Verification
Cases in benchmarks/agent/test_agent_bench.py; one real run recorded in the PR as evidence that the level works end to end (not as a product claim).
Dependencies / ownership
Owner: Wright (benchmarks/agent). Parent: #470. Depends on the MCP adapter Issue. Orthogonal to #466.
Problem
There is no measurement of whether native Wright tools improve on shell + JSON for task success and search/read/tool-call efficiency. See ADR-0020 (Proposed in the ADR PR for #470).
Scope
Add
mcpas awrightlevel besidebinin the agent benchmark condition grid (<wright>/<knowledge>/<network>), with scenario, prompt, grader, model, and trial identity unchanged:wright serve --transport mcpserver for the agent;wrightUse,friction, and the report cover them;mcpwithbinon usable/passed rate with Wilson intervals and on search/read/tool-call counts, turns, and tokens; tool-schema context counts toward themcpcost;docs/agent-benchmark.mddocuments the level.Non-goals
Acceptance criteria
agent_bench.py validateand the existing benchmark unit tests pass unchanged.mcpproducesresult.jsonwith MCP calls counted inwrightUse, and a canary fails the run asinvalidwhenwrightis reachable under levelnone, as today.reportshowsmcpversusbinpaired on the same scenario, agent, and trial, with the metrics listed above.mcpinput-token accounting, verified on a fixture usage log.Verification
Cases in
benchmarks/agent/test_agent_bench.py; one real run recorded in the PR as evidence that the level works end to end (not as a product claim).Dependencies / ownership
Owner: Wright (
benchmarks/agent). Parent: #470. Depends on the MCP adapter Issue. Orthogonal to #466.