Skip to content

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Yanuo Ma*, Ben Kereopa-Yorke, Ben Schultz (Microsoft)
arXiv 2606.28430 · Local PDF: paper/arxiv.pdf

Artifacts for the paper above — the experimental corpus, audit tool, methodology, and task spec used to produce it. The paper studies whether two production LLM coding agents — GitHub Copilot CLI with claude-opus-4.7 and gpt-5.5 — deliver the artifact they were asked to build, when scored by an honest hidden behavioral oracle.

What's here

Three categories, distinguished by location and naming convention:

  1. Pre-authored harness — everything not under runs/ and not prefixed with _. The inputs given to the agents and the orchestration around them.
  2. Agent experiment data — runs/<agent>/fluent-table/<cond>/<run-id>/ minus anything prefixed with _. The 18 per-run workspaces and session logs the agents produced.
  3. Post-hoc analysis — anything prefixed with _, at any depth. Authored by the paper team after the experiment for scoring or analysis.
Path What Category Cited from
paper/arxiv.pdf The preprint — —
prompts/ Verbatim agent-facing files (TASK.md, AGENTS-c{0,3,9}.md, fetch.sh, oracle-allpass-sample.txt) Pre-authored App A.1–A.4, B, C
tasks/ui/fluent-table/ Task specification (the agent's full input space) Pre-authored App A, B
scripts/copilot/ Orchestration scripts (Docker, Copilot CLI) Pre-authored —
runs/<agent>/fluent-table/<cond>/<run-id>/ 18 per-run workspaces and session logs Agent data Apps F, J, K
_audit-tool/ Static demo-reference audit (the Table 1 verdicts) Post-hoc App E
runs/_ablation-evidence/ No-op ablation diffs and outcomes Post-hoc App H
runs/<agent>/c{0,c3-R3}/.../workspace/_eval-storybook/ Consumer kits authored post-hoc to score c0 deliverables (and GPT c3-R3) Post-hoc App G

Naming notes

  • wild, wild-runs, wild-test — internal project names from the methodology's early development. They appear pervasively in agent AGENTS.md files (the helper tool the agent invokes is literally named wild-test), session logs, and orchestration scripts. They carry no specific meaning; readers can mentally substitute b2t. Renaming them post-experiment would misrepresent what the agent actually saw.
  • claude-opus-4.7-1m-internal — Microsoft-internal label for the 1M-context variant of Claude Opus 4.7; preserved as-is in the corpus and the paper for accuracy.
  • <repo-root>, <wild-repo>, <host-home>, <oracle-test>, <oracle-harness> — placeholders introduced by path scrubbing in post-hoc analysis files (oracle JSON outputs, ablation diffs, validation notes). The original paths were on the author's host machine and added nothing reproducible.

Reproducing the study

Prerequisites (one-time, host)

  • Docker, current user in docker group (docker run --rm hello-world succeeds without sudo)
  • GitHub Copilot CLI installed and authed on the host (~/.copilot/config.json present)
  • jq

Step 1 — Build the runner image

./scripts/copilot/build-image.sh

Produces wild-ui:latest (~3.4 GB, ~5–10 min first build). The image bakes:

  • Node 24 + pnpm + @github/copilot (pinned in tasks/ui/docker/Dockerfile)
  • Chromium + Playwright system libraries
  • The 222-test oracle at /opt/wild-tests/playwright/ (root:root, mode 700 — the runner user cannot read the specs)
  • The wild-test setuid wrapper — the only path by which the agent can trigger the oracle without reading the source

Step 2 — Start a single run

./scripts/copilot/start-run.sh ui/fluent-table \
    --oracle c3 \
    --auto --model claude-opus-4.7 --effort xhigh

Flags:

  • --oracle {c0,c3,c9} — oracle-availability condition (§4)
    • c0 — no wild-test, oracle unavailable to the agent
    • c3 — wild-test available under the guardrail framing in prompts/AGENTS-c3.md (helper, not goal)
    • c9 — wild-test available under the deliverable framing in prompts/AGENTS-c9.md (no hedge)
  • --auto --model <name> --effort <level> — non-interactive mode; values used in the paper are claude-opus-4.7 / gpt-5.5, effort xhigh

Creates the working tree at ~/wild-runs/working/<run-id>/, copies your Copilot auth into the container (whitelist-copied, host ~/.copilot/ is never bind-mounted), and drops you into docker start -ai. Type copilot inside to begin. Exit the shell when the agent stops.

Step 3 — Export the artifacts and clean up

./scripts/copilot/export-run.sh  <run-id>
./scripts/copilot/cleanup-run.sh <run-id>

export-run.sh writes the workspace + session log into a per-run archive that matches the layout under runs/ in this repo.

Step 4 — Batch (all 18 runs)

./scripts/copilot/batch-run.sh

Iterates the (model × condition × replicate) matrix used in the paper. See the script header for tunable parameters.

Step 5 — Score the oracle against delivered code

For c3/c9 runs the agent triggered the oracle in-loop via wild-test; results are in the exported run's session/events.jsonl. For c0 runs the agent never invoked the oracle by design (that is the c0 condition), so scoring the delivery is a post-hoc pass. Either way, run the tests standalone against the agent's workspace:

cd tasks/ui/fluent-table/harness/playwright
pnpm install && npx playwright install chromium
BASE_URL=http://localhost:6006 npx playwright test

BASE_URL points to a Storybook that exposes the expected story IDs at /iframe.html?id=<storyId> — typically the agent's own workspace, or a consumer kit built to stand in for the delivered library.

Verifying our released numbers

_audit-tool/ and runs/_ablation-evidence/ are post-hoc analyses hardcoded to this specific 12-run c3/c9 corpus. python3 _audit-tool/audit.py --workspaces-root ./runs --verify reports 48/48 agreement against _audit-tool/expected.json (Table 1); the ablation dossier is static and reviewed directly. Applying the same methodology to a different task or agent requires authoring an equivalent audit tool and ablation matrix for that task's expected library shape.

Appendix K trajectory quotes

Cited verbatim from runs/<agent>/fluent-table/<cond>/<run-id>/session/events.jsonl (field data.reasoningText or data.content).

Applying this methodology elsewhere

The paper's contribution is a method — code-as-spec + strict oracle-exposure control + post-hoc audit — not a specific test suite. The same recipe applies to any task where a runnable reference exists:

  • UI component libraries (any framework pair: Material UI → Vue, Mantine → Svelte, ag-Grid → any)
  • CLI tools with behavioral test suites
  • APIs with consumer-driven contract tests
  • Domain-specific languages with reference interpreters

A future replication on a new task with a different oracle is methodologically stronger than reusing our specific 222 tests — it eliminates any concern about our oracle's particular shape biasing the results.

Repository status

This is the as-released snapshot accompanying the arXiv preprint. It is a single-shot release rather than an evolving codebase; subsequent work using the same methodology will live in separate repositories.

License

MIT — see LICENSE. This matches the Fluent UI React reference's own license. The audit tool and methodology are released for any use, academic or commercial.

Citation

@article{ma2026buildingtothetest,
  title     = {Building to the Test: Coding Agents Deliver What You Check, Not What You Requested},
  author    = {Ma, Yanuo and Kereopa-Yorke, Ben and Schultz, Ben},
  journal   = {arXiv preprint arXiv:2606.28430},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.28430}
}

Contact

Correspondence: yanuoma@microsoft.com

About

Artifacts for arXiv:2606.28430. Task spec, prompts, 18-run agent corpus, and a deterministic audit tool from a study showing two production LLM coding agents (Copilot CLI · claude-opus-4.7, gpt-5.5) score near-perfect on a hidden 222-test oracle while leaving the requested library dead or absent.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages