test: add scripted-inference e2e layer and slim the unit suite - #1221
Merged
TheGreatAxios merged 35 commits intoSep 29, 2026
Merged
TheGreatAxios merged 35 commits into
TheGreatAxios merged 35 commits into
Conversation
Persisted tool_call blocks keep the wire name the model emitted (read, glob, bash), but the compactor's replay/supersession index and the subagent thrash tracker classify against engine-keyed sets (read_file, search_files). After the one-wire migration that left read supersession and thrash bookkeeping dead for the advertised tools; route stored names through canonicalToolName before classification.
The e2e layer wraps the existing agent-loop harness (live tool dispatch, chat director, posix tools, git-backed context store) with fixture-repo seeding and a sendOperatorTurn driver that surfaces terminal failures instead of asserting a reply. The integration session gains the seams scenarios need: fleet passthroughs, retry-policy and scheduler overrides, alternate inference sources plus credential records, and access to the live chat director. The toolset forwards the fleet's documented retry-timing overrides to spawned runs.
Three orchestration scenarios now run through the real agent loop on scripted inference: compaction read-paging (the wire-name coverage that exposed the engine-name classification bug), recoverable subagent failure (failed lane reports continuable through wait_agents), and credential recovery (oauth refresh, retry, terminal 401, accepted backup serves the replayed prompt). The bespoke unit versions are deleted where the e2e covers the same contract; constant-equal-literal tool-alias pins and a decisions-diagnostics-only tail test are dropped for asserting the same thing behavior already proves. AGENTS.md documents the split and the v1 non-goals.
The engine-name fix for persisted calls had no fixture proving wire names reach the same evidence keys — regression coverage for the read/search/mutation path that silently went dark after the one-wire migration.
A stream error in the background collector would surface as an unhandled rejection in the shared test process; the catch mirrors runUntilSuspended. Explicit timeouts replace the 5s bun:test default so a loaded CI worker fails a genuinely stuck scenario rather than flaking at the default.
Cuts ~10.6k lines of co-located unit tests (~7%) plus ~1.4k of assertion-level copy/literal/diagnostic pins, and folds ~3.4k lines of copy-pasted arrange scaffolding into file-local helpers with zero assertion loss. Moves the ask_director handshake and outer-retry recovery scenarios into tests/e2e driven through the real agent loop.
TheGreatAxios
force-pushed
the
cl-9468-add-scripted-inference-e2e-layer-migrate-orchestration-tests
branch
from
September 29, 2026 17:29
9d4e06e to
c5dd317
Compare
TheGreatAxios
deleted the
cl-9468-add-scripted-inference-e2e-layer-migrate-orchestration-tests
branch
September 29, 2026 17:32
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
e2e/layer: scripted-inference scenario tests (e2e/harness.ts,e2e/fleet.ts) that drive the production agent/reactor/director/tool loop over fixture repos via@intx/inference-testing, plus the formertests/integrationsuites relocated theresrc/**/*.test.ts), moves shared helpers tosrc/testkit/and fixture repos tofixtures/; thetests/directory is removedfleet-test-harness,run-test-harness,withAppShell,reactor-stubs) — net ~−24k test lines and ~−1.1k testsisSessionMode,isOtelConfigInvalid, orphaned exportsVerification
bun run typecheck,bun run build, andbun run testpassbun run checkgreen: oxlint/oxfmt, dead-export guard (0 violations), projects-dir-guarded randomized suite — 7,541 pass / 0 fail / 509 filesFixes CL-9468