Yanuo Ma*, Ben Kereopa-Yorke, Ben Schultz (Microsoft)
arXiv 2606.28430 · Local PDF:paper/arxiv.pdf
Artifacts for the paper above — the experimental corpus, audit tool,
methodology, and task spec used to produce it. The paper studies whether two
production LLM coding agents — GitHub Copilot CLI with claude-opus-4.7 and
gpt-5.5 — deliver the artifact they were asked to build, when scored by an
honest hidden behavioral oracle.
Three categories, distinguished by location and naming convention:
- Pre-authored harness — everything not under
runs/and not prefixed with_. The inputs given to the agents and the orchestration around them. - Agent experiment data —
runs/<agent>/fluent-table/<cond>/<run-id>/minus anything prefixed with_. The 18 per-run workspaces and session logs the agents produced. - Post-hoc analysis — anything prefixed with
_, at any depth. Authored by the paper team after the experiment for scoring or analysis.
| Path | What | Category | Cited from |
|---|---|---|---|
paper/arxiv.pdf |
The preprint | — | — |
prompts/ |
Verbatim agent-facing files (TASK.md, AGENTS-c{0,3,9}.md, fetch.sh, oracle-allpass-sample.txt) |
Pre-authored | App A.1–A.4, B, C |
tasks/ui/fluent-table/ |
Task specification (the agent's full input space) | Pre-authored | App A, B |
scripts/copilot/ |
Orchestration scripts (Docker, Copilot CLI) | Pre-authored | — |
runs/<agent>/fluent-table/<cond>/<run-id>/ |
18 per-run workspaces and session logs | Agent data | Apps F, J, K |
_audit-tool/ |
Static demo-reference audit (the Table 1 verdicts) | Post-hoc | App E |
runs/_ablation-evidence/ |
No-op ablation diffs and outcomes | Post-hoc | App H |
runs/<agent>/c{0,c3-R3}/.../workspace/_eval-storybook/ |
Consumer kits authored post-hoc to score c0 deliverables (and GPT c3-R3) | Post-hoc | App G |
wild,wild-runs,wild-test— internal project names from the methodology's early development. They appear pervasively in agentAGENTS.mdfiles (the helper tool the agent invokes is literally namedwild-test), session logs, and orchestration scripts. They carry no specific meaning; readers can mentally substituteb2t. Renaming them post-experiment would misrepresent what the agent actually saw.claude-opus-4.7-1m-internal— Microsoft-internal label for the 1M-context variant of Claude Opus 4.7; preserved as-is in the corpus and the paper for accuracy.<repo-root>,<wild-repo>,<host-home>,<oracle-test>,<oracle-harness>— placeholders introduced by path scrubbing in post-hoc analysis files (oracle JSON outputs, ablation diffs, validation notes). The original paths were on the author's host machine and added nothing reproducible.
- Docker, current user in
dockergroup (docker run --rm hello-worldsucceeds withoutsudo) - GitHub Copilot CLI installed and authed on the host (
~/.copilot/config.jsonpresent) jq
./scripts/copilot/build-image.shProduces wild-ui:latest (~3.4 GB, ~5–10 min first build). The image bakes:
- Node 24 + pnpm +
@github/copilot(pinned intasks/ui/docker/Dockerfile) - Chromium + Playwright system libraries
- The 222-test oracle at
/opt/wild-tests/playwright/(root:root, mode 700 — the runner user cannot read the specs) - The
wild-testsetuid wrapper — the only path by which the agent can trigger the oracle without reading the source
./scripts/copilot/start-run.sh ui/fluent-table \
--oracle c3 \
--auto --model claude-opus-4.7 --effort xhighFlags:
--oracle {c0,c3,c9}— oracle-availability condition (§4)c0— nowild-test, oracle unavailable to the agentc3—wild-testavailable under the guardrail framing inprompts/AGENTS-c3.md(helper, not goal)c9—wild-testavailable under the deliverable framing inprompts/AGENTS-c9.md(no hedge)
--auto --model <name> --effort <level>— non-interactive mode; values used in the paper areclaude-opus-4.7/gpt-5.5, effortxhigh
Creates the working tree at ~/wild-runs/working/<run-id>/, copies your
Copilot auth into the container (whitelist-copied, host ~/.copilot/ is
never bind-mounted), and drops you into docker start -ai. Type copilot
inside to begin. Exit the shell when the agent stops.
./scripts/copilot/export-run.sh <run-id>
./scripts/copilot/cleanup-run.sh <run-id>export-run.sh writes the workspace + session log into a per-run archive that
matches the layout under runs/ in this repo.
./scripts/copilot/batch-run.shIterates the (model × condition × replicate) matrix used in the paper. See the script header for tunable parameters.
For c3/c9 runs the agent triggered the oracle in-loop via wild-test;
results are in the exported run's session/events.jsonl. For c0 runs the
agent never invoked the oracle by design (that is the c0 condition), so
scoring the delivery is a post-hoc pass. Either way, run the tests
standalone against the agent's workspace:
cd tasks/ui/fluent-table/harness/playwright
pnpm install && npx playwright install chromium
BASE_URL=http://localhost:6006 npx playwright testBASE_URL points to a Storybook that exposes the expected story IDs at
/iframe.html?id=<storyId> — typically the agent's own workspace, or a
consumer kit built to stand in for the delivered library.
_audit-tool/ and runs/_ablation-evidence/
are post-hoc analyses hardcoded to this specific 12-run c3/c9 corpus.
python3 _audit-tool/audit.py --workspaces-root ./runs --verify reports
48/48 agreement against _audit-tool/expected.json (Table 1); the ablation
dossier is static and reviewed directly. Applying the same methodology to a
different task or agent requires authoring an equivalent audit tool and
ablation matrix for that task's expected library shape.
Cited verbatim from
runs/<agent>/fluent-table/<cond>/<run-id>/session/events.jsonl (field
data.reasoningText or data.content).
The paper's contribution is a method — code-as-spec + strict oracle-exposure control + post-hoc audit — not a specific test suite. The same recipe applies to any task where a runnable reference exists:
- UI component libraries (any framework pair: Material UI → Vue, Mantine → Svelte, ag-Grid → any)
- CLI tools with behavioral test suites
- APIs with consumer-driven contract tests
- Domain-specific languages with reference interpreters
A future replication on a new task with a different oracle is methodologically stronger than reusing our specific 222 tests — it eliminates any concern about our oracle's particular shape biasing the results.
This is the as-released snapshot accompanying the arXiv preprint. It is a single-shot release rather than an evolving codebase; subsequent work using the same methodology will live in separate repositories.
MIT — see LICENSE. This matches the Fluent UI React reference's own license. The audit tool and methodology are released for any use, academic or commercial.
@article{ma2026buildingtothetest,
title = {Building to the Test: Coding Agents Deliver What You Check, Not What You Requested},
author = {Ma, Yanuo and Kereopa-Yorke, Ben and Schultz, Ben},
journal = {arXiv preprint arXiv:2606.28430},
year = {2026},
url = {https://arxiv.org/abs/2606.28430}
}Correspondence: yanuoma@microsoft.com