Skip to content

test: reduce browser journey fixture and shard setup - #302

Closed
kalvinnchau wants to merge 1 commit into
mainfrom
coffee/browser-ci-setup
Closed

kalvinnchau wants to merge 1 commit into
mainfrom
coffee/browser-ci-setup

Conversation

@kalvinnchau

@kalvinnchau kalvinnchau commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Goal

Reduce repeated browser-journey setup without reducing coverage, assertions, engine support, or measurement isolation.

Changes

  • Reuse the existing worker-scoped source server for agent-control and messages journeys. Add one worker-scoped sourcePlugins option for messages' stateless media middleware. Browser contexts, storage and fixture identities remain test-local.
  • Provision only each functional shard's selected engine; the isolated measurement job still provisions Chromium and WebKit. Browser cache keys include the engine selection.
  • Discover each shard with the same Playwright selection used for execution. Only selections containing the conversation journey's @native-fixture tag restore the Rust caches/build fixture-bridge. Empty/failed discovery fails closed; cache misses still build. No shard number is hard-coded.
  • Keep the selected Rust toolchain independently cached: non-native shards may finish first and save the shared Hermit cache before Rust is present.
  • Extend the workflow guards and add a focused real-Playwright lifecycle control.

No application code, retries, timeouts, worker/shard counts, measurement selection, or existing assertions changed. Most of the large textual diff is dedenting test bodies after removing their per-test server try/finally wrappers. An AST comparison verified all 264 agent-control and 224 messages expect call nodes are unchanged.

Local before/after evidence

Environment: Apple Silicon macOS, pinned Node 24.18.0 / Playwright 1.63.0, two workers. One sample per engine, not a statistical or hosted-CI speedup claim. Dependencies/browser engines were already installed. Suite wall time includes browser/fixture setup; individual Vite setup time was not instrumented separately.

Command (each engine separately):

PLAYWRIGHT_JSON_OUTPUT_FILE=<report> bin/pnpm test:browser:ci \
  tests/browser/agent-control.spec.mjs tests/browser/messages.spec.mjs \
  --project <chromium|webkit> --no-deps --workers=2 --reporter=list,json

Before: clean 64be4c2a. After: the same base with this PR's fixture changes. The branch subsequently incorporated 48c9a11d; the affected files plus its profile journeys were rerun successfully.

Engine Cases before/after Wall before → after Summed execution before → after
Chromium 20/20, all passed 41.63s → 34.16s 78.46s → 59.11s
WebKit 20/20, all passed 51.18s → 37.54s 94.65s → 70.47s
File Chromium summed before → after WebKit summed before → after
agent-control 40.12s → 25.89s 50.20s → 33.77s
messages 38.34s → 33.23s 44.45s → 36.70s

The slowest case remained the shared-thread UI journey: Chromium 16.52s → 16.68s, WebKit 18.98s → 17.16s. This change removes repeated setup rather than shortening that journey's assertions.

Hosted baseline: successful main run 36179555533, 4012979b, Ubuntu 24.04 x64: slowest WebKit shard 11m50s total / 10m21s functional command; shared setup 50s, Rust cache 5s, native build 28s. All six shards previously built the fixture (20–28s each). Different source snapshots and cold/new cache keys prevent treating a single PR run as a controlled speedup comparison. Hosted evidence is recorded below.

Regression coverage and validation

  • bin/node --test tests/integration/browser-ci.test.mjs tests/integration/vite-fixture.test.mjs: 11 passed. Executes the actual discovery shell for all six shards, verifies complete nonduplicated selection, native setup ownership, exact engine-install arguments, invalid input rejection, cache guards and the required gate.
  • Existing 20 affected browser cases retained in both engines: 0 removed or moved. No replacement assertions needed.
  • 2 browser controls added (4 engine cases): actual Playwright worker reuse, custom static middleware and fresh per-test localStorage/cookies. A Node-only test cannot establish the runner/browser-context boundary. Mutating the server fixture back to test scope fails all four controls (starts becomes 2 instead of 1); restoring worker scope passes. The controls also passed three repetitions in both engines.
  • Full affected files plus lifecycle controls, both engines, --repeat-each=2: 88 passed, zero retries/skips. An initial static-control-only optimizer cleanup error was fixed by avoiding unrelated dependency discovery in that HTML-only control; product fixture optimization is unchanged.
  • After incorporating incoming profile changes: agent-control, messages, source-fixture, profiles and profile-instances files in both engines: 56 passed on the tree committed as 06ce001f.
  • Biome checks and git diff --check passed. Existing organization hooks and repository hooks both ran; pre-push correctly selected no source-unit/design inputs for this tooling/test-only change.
  • Self-review: no known material correctness defects; checked selection parity, failure behavior, cache misses, isolation, cleanup and unchanged assertions. Independent review remains pending.

Hosted result at 06ce001f

CI run 36187164665 passed, including the required aggregate, JavaScript, Rust/tool integration and both browser lanes. DCO, Semgrep and zizmor also passed. Browser artifacts identify the checked PR merge tree as 5d1a9fb9.

  • 716/716 functional cases passed (358 per engine); 7/7 measurement cases passed; no skips or retries. The increase from 712 is exactly the four engine instances of the new lifecycle controls.
  • Only Chromium 1/3 and WebKit 1/3 selected the tagged native journey on this snapshot. Their native builds took 27s/25s; all Rust setup/cache steps were skipped on the other four runners. This assignment is discovered, not encoded.
  • Discovery cost 2–5s per shard. Shared setup cost 48–84s; the new engine-specific browser cache keys and the Rust-only toolchain cache were cold. Logs confirm WebKit runners downloaded WebKit, not Chromium, and vice versa. Measurements still provisioned both.

Nearest passing baseline: 36185070394, 64be4c2a, the same Ubuntu 24.04 x64 / Playwright 1.63.0 environment. This is an observational comparison, not a controlled benchmark: the PR merge tree includes later main changes, fixture ownership changes file grouping, runner speed varies, and cache warmth differs.

Job Baseline whole-job wall PR whole-job wall PR functional command
Chromium 1/3 10m18s 9m20s 7m34s
Chromium 2/3 7m01s 4m52s 3m45s
Chromium 3/3 7m25s 8m02s 6m57s
WebKit 1/3 8m03s 11m00s 8m46s
WebKit 2/3 8m54s 7m38s 6m08s
WebKit 3/3 10m58s 10m16s 9m11s

Summed test execution: Chromium 2165.50s → 1997.03s; WebKit 2592.40s → 2646.06s. Measurements: 170.85s → 169.35s command wall, 159.62s → 158.88s summed execution.

Changed-file execution sums: agent-control Chromium 134.4s → 75.1s, WebKit 90.9s → 90.9s; messages Chromium 116.6s → 80.9s, WebKit 65.5s → 84.8s. The remaining slowest WebKit files are nested-replies (226.3s) and message-navigation (157.9s). The slowest job remains about 11 minutes: this run does not establish a whole-suite speedup. Repeated, same-snapshot warm-cache comparisons remain deferred rather than claiming the local reductions transfer directly to CI.

Remaining checks / human verification

  • Hosted CI, DCO and post-change timing/step evidence at the PR head.
  • Repeated same-snapshot warm-cache timing comparisons before making a whole-suite speedup claim.
  • Independent review.
  • Human confirmation: inspect the browser matrix to confirm every shard succeeds, only shards selecting the native journey run its Rust setup, and the measurement lane retains both engines. Optionally run the focused command above in both engines; all 20 original cases should pass.

Draft until those checks and human confirmation are complete.

Signed-off-by: Kalvin Chau <kalvin@block.xyz>
@kalvinnchau

Copy link
Copy Markdown
Contributor Author

Self-review of 48c9a11d..06ce001f: no blocking findings after checking selection parity, failed/empty discovery, cache-miss provisioning, worker/context isolation, teardown and unchanged existing assertions. The real discovery guard caught the JSON tag normalization issue during development; the final implementation passes all six selection checks. Mutating source-server scope back to per-test makes all four lifecycle controls fail; the restored implementation passes. Independent review and human confirmation remain pending. CI/DCO are green; the PR body records counts, setup/execution timings and the important limitation that the slowest hosted job is still approximately 11 minutes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant