Skip to content

test: add scripted-inference e2e layer and slim the unit suite - #1221

Merged
TheGreatAxios merged 35 commits into
mainfrom
cl-9468-add-scripted-inference-e2e-layer-migrate-orchestration-tests
Sep 29, 2026
Merged

TheGreatAxios merged 35 commits into
mainfrom
cl-9468-add-scripted-inference-e2e-layer-migrate-orchestration-tests

Conversation

@TheGreatAxios

Copy link
Copy Markdown
Collaborator

Summary

  • Adds a root-level e2e/ layer: scripted-inference scenario tests (e2e/harness.ts, e2e/fleet.ts) that drive the production agent/reactor/director/tool loop over fixture repos via @intx/inference-testing, plus the former tests/integration suites relocated there
  • Co-locates unit tests beside their modules (src/**/*.test.ts), moves shared helpers to src/testkit/ and fixture repos to fixtures/; the tests/ directory is removed
  • Deletes low-value unit tests (literal copy pins, constant self-pins, mock tautologies, strict duplicates, unit replicas of e2e-covered flows) and collapses repeated scaffolding into shared harnesses (fleet-test-harness, run-test-harness, withAppShell, reactor-stubs) — net ~−24k test lines and ~−1.1k tests
  • Removes dead production code surfaced by the prune: write-only Codex-404 telemetry, unused prompt-note builders, isSessionMode, isOtelConfigInvalid, orphaned exports
  • Fixes a real bug the e2e layer caught: persisted tool calls were classified by display name instead of wire name, breaking compaction tracking for aliased tools

Verification

  • bun run typecheck, bun run build, and bun run test pass
  • bun run check green: oxlint/oxfmt, dead-export guard (0 violations), projects-dir-guarded randomized suite — 7,541 pass / 0 fail / 509 files

Fixes CL-9468

@linear-code

linear-code Bot commented Sep 29, 2026

Copy link
Copy Markdown

CL-9468

Persisted tool_call blocks keep the wire name the model emitted
(read, glob, bash), but the compactor's replay/supersession index
and the subagent thrash tracker classify against engine-keyed sets
(read_file, search_files). After the one-wire migration that left
read supersession and thrash bookkeeping dead for the advertised
tools; route stored names through canonicalToolName before
classification.
The e2e layer wraps the existing agent-loop harness (live tool
dispatch, chat director, posix tools, git-backed context store)
with fixture-repo seeding and a sendOperatorTurn driver that
surfaces terminal failures instead of asserting a reply. The
integration session gains the seams scenarios need: fleet
passthroughs, retry-policy and scheduler overrides, alternate
inference sources plus credential records, and access to the live
chat director. The toolset forwards the fleet's documented
retry-timing overrides to spawned runs.
Three orchestration scenarios now run through the real agent loop
on scripted inference: compaction read-paging (the wire-name
coverage that exposed the engine-name classification bug),
recoverable subagent failure (failed lane reports continuable
through wait_agents), and credential recovery (oauth refresh,
retry, terminal 401, accepted backup serves the replayed prompt).
The bespoke unit versions are deleted where the e2e covers the
same contract; constant-equal-literal tool-alias pins and a
decisions-diagnostics-only tail test are dropped for asserting
the same thing behavior already proves. AGENTS.md documents the
split and the v1 non-goals.
The engine-name fix for persisted calls had no fixture proving
wire names reach the same evidence keys — regression coverage for
the read/search/mutation path that silently went dark after the
one-wire migration.
A stream error in the background collector would surface as an
unhandled rejection in the shared test process; the catch mirrors
runUntilSuspended. Explicit timeouts replace the 5s bun:test
default so a loaded CI worker fails a genuinely stuck scenario
rather than flaking at the default.
Cuts ~10.6k lines of co-located unit tests (~7%) plus ~1.4k of
assertion-level copy/literal/diagnostic pins, and folds ~3.4k lines of
copy-pasted arrange scaffolding into file-local helpers with zero
assertion loss. Moves the ask_director handshake and outer-retry
recovery scenarios into tests/e2e driven through the real agent loop.
@TheGreatAxios
TheGreatAxios force-pushed the cl-9468-add-scripted-inference-e2e-layer-migrate-orchestration-tests branch from 9d4e06e to c5dd317 Compare September 29, 2026 17:29
@TheGreatAxios
TheGreatAxios merged commit 065075f into main Sep 29, 2026
13 checks passed
@TheGreatAxios
TheGreatAxios deleted the cl-9468-add-scripted-inference-e2e-layer-migrate-orchestration-tests branch September 29, 2026 17:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant