[ci] Cap concurrent Vercel E2E lanes repo-wide - #4039
Conversation
The Vercel E2E lanes in tests.yml fan out to 38 jobs per run (27 in e2e-vercel-prod, 6 HTTP transport, 4 WS transport, 1 multi-region), and every one of them drives real traffic at world-vercel and at a shared Vercel project. Nothing bounded that across runs: the only concurrency control in the file is the workflow-level group, which dedupes a branch against itself but lets N open PRs multiply the load by N. Give each Vercel lane a job-level concurrency group keyed on (lane, matrix cell, slot). Within one run every cell is a distinct group, so a single PR still fans out in full; the same cell from a different run queues instead. The slot count, computed in ci-scope because GitHub expressions have no arithmetic, sets how many runs' worth may be in flight repo-wide and is tunable via the E2E_VERCEL_CONCURRENCY_SLOTS repository variable (default 2). `queue: max` is required rather than incidental. The default `queue: single` holds one pending entry per group and cancels the rest, which would convert a burst of PRs into cancelled jobs and a red E2E Required Check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 3e49a51 The changes in this PR will be included in the next version bump. This PR includes changesets to release 0 packagesWhen changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ✅ ▲ Vercel Production | 3662 | 0 | 685 | 4347 |
| ✅ 💻 Local Development | 3922 | 0 | 586 | 4508 |
| ✅ 📦 Local Production | 3922 | 0 | 586 | 4508 |
| ✅ 🐘 Local Postgres | 3922 | 0 | 586 | 4508 |
| ✅ 🪟 Windows | 320 | 0 | 2 | 322 |
| ✅ 🌐 Cross-language Conformance | 68 | 0 | 74 | 142 |
| ✅ vercel-http-transport | 823 | 0 | 143 | 966 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 557 | 0 | 87 | 644 |
| Total | 17223 | 0 | 2749 | 19972 |
Details by Category
✅ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 133 | 0 | 28 |
| ✅ astro-quickjs | 133 | 0 | 28 |
| ✅ example-node | 133 | 0 | 28 |
| ✅ example-quickjs | 133 | 0 | 28 |
| ✅ express-node | 133 | 0 | 28 |
| ✅ express-quickjs | 133 | 0 | 28 |
| ✅ fastify-node | 133 | 0 | 28 |
| ✅ fastify-quickjs | 133 | 0 | 28 |
| ✅ hono-node | 133 | 0 | 28 |
| ✅ hono-quickjs | 133 | 0 | 28 |
| ✅ nest-node | 133 | 0 | 28 |
| ✅ nest-quickjs | 133 | 0 | 28 |
| ✅ nextjs-turbopack-node | 158 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 158 | 0 | 3 |
| ✅ nextjs-webpack-node | 158 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 158 | 0 | 3 |
| ✅ nitro-node | 133 | 0 | 28 |
| ✅ nitro-quickjs | 133 | 0 | 28 |
| ✅ nuxt-node | 133 | 0 | 28 |
| ✅ nuxt-quickjs | 133 | 0 | 28 |
| ✅ python-node | 66 | 0 | 95 |
| ✅ sveltekit-node | 152 | 0 | 9 |
| ✅ sveltekit-quickjs | 152 | 0 | 9 |
| ✅ tanstack-start-node | 133 | 0 | 28 |
| ✅ tanstack-start-quickjs | 133 | 0 | 28 |
| ✅ vite-node | 133 | 0 | 28 |
| ✅ vite-quickjs | 133 | 0 | 28 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 160 | 0 | 1 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 68 | 0 | 74 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 28 |
| ✅ express | 133 | 0 | 28 |
| ✅ hono | 133 | 0 | 28 |
| ✅ nextjs-turbopack | 158 | 0 | 3 |
| ✅ nitro | 133 | 0 | 28 |
| ✅ vite | 133 | 0 | 28 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 28 |
| ✅ express | 133 | 0 | 28 |
| ✅ nextjs-turbopack | 158 | 0 | 3 |
| ✅ vite | 133 | 0 | 28 |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 120942ms → this run 135981ms (Δ +15039ms, +12%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
|
No backport to This adds a new CI capacity-management capability (per-cell/slot job concurrency groups plus a tunable To override, re-run the Backport to stable workflow manually via |
…c-workflow-source * origin/main: (22 commits) feat(streams): add writer session seam (#3832) docs: document WORKFLOW_NODE_HTTP in the v4 World docs (#4050) [world-vercel] Honor WORKFLOW_NODE_HTTP on the queue transport (#4044) feat(streams): add WebSocket capability gate (#3764) fix(world-local): retry JSON reads on Windows (#4051) [core] Add the wake-loop scenario to the event log race repro (#4017) [ci] Cap concurrent Vercel E2E action repo-wide (10 by default) (#4039) Add `Run#getWritable()` for appending to another run's stream (#3972) feat(streams): add WebSocket v1 client protocol contract (#3763) [swc-playground] Update to Next.js v16.3.4 (#4018) [core] Add a retention option to start() (#3787) fix(swc-plugin): register class expressions via an IIFE instead of by name (#3971) [docs] Fix prose typos across v4/v5 docs and the SWC plugin README (#3948) Validate pending changesets in CI so a bad one fails the PR, not the Release job (#3964) Version Packages (beta) (#3919) Classify Workflow stream failures (#3850) Drop the ignored @workflow/world-sim package from the hook_conflict delta changeset (#3963) Add attribute inspection to the CLI (#3950) [core] Settle a hook's awaiter in-process instead of re-invoking, on creation and on conflict (#3938) Use the storage APIs for run detail views (#3944) ...
Problem
The Vercel E2E lanes in
tests.ymlfan out to 38 jobs per run (27e2e-vercel-prod, 6 HTTP transport, 4 WS transport, 1 multi-region). Each drives real traffic atapi-workflow, and nothing bounded that across runs: the only concurrency control in the file dedupes a branch against itself but lets N open PRs multiply load by N.That the E2E runner is an
api-workflowclient is a property of the code path, not an inference.packages/world-vercel/src/utils.ts:CI sets
WORKFLOW_VERCEL_PROJECT+WORKFLOW_VERCEL_TEAM+WORKFLOW_VERCEL_AUTH_TOKENon every Vercel lane, so the test runner takes the proxy. The deployed app under test (OIDC, noprojectConfig) goes direct.Does GitHub support a cap?
Partly, and only since May 2026. There is still no "max N concurrent" setting: a concurrency group admits one running entry. What changed is the queue behind it.
queue: maxlets up to 100 entries wait instead of the defaultqueue: single, which keeps one pending entry and cancels the rest. Without it, capping a required check would turn a burst of PRs into cancelled jobs and a red E2E Required Check. So the cap is built from many one-at-a-time groups rather than one N-at-a-time group.queueis documented under workflow-levelconcurrencyand not shown in thejobs.<job_id>.concurrencysection, so I probed it: one job, one fixed group,queue: max, three pushes. Two runs satpendingin the group simultaneously (impossible underqueue: single), none was cancelled, and they executed strictly FIFO. Job-levelqueue: maxis honored.Sizing, from INC-7881
Rather than pick a number, I correlated this file's Vercel lanes against
cluster_monitor_production.hpa_desired_replicas{hpa_name:api-workflow,env:sfo1}for 2026-09-08 19:00–22:30 UTC, per minute:Pearson r = 0.43.
Two things follow, and the first is a caveat on this PR. The HPA reaches its ceiling 41% of the time with CI completely idle, and the earlier 33-minute saturation at 19:16–19:49 (which also crossed the monitor's 30-minute window) began with zero e2e jobs running. CI is an aggravating load, not the whole story, and this change alone will not keep
api-workflowoff its ceiling. Second: CI's own contribution stops being marginal somewhere above 150 concurrent jobs, which is what sizes the tiers.What the incident peak was made of
At 20:59, the minute the HPA pinned, all 379 concurrent Vercel e2e jobs came from ten
test/server-pr-*branches — the companion PRs aworkflow-serverPR opens here — each running the full 38-cell fan-out. Not one was a human PR.So the cap is tiered. Machine-opened branches (
test/server-pr-*,changeset-release/*) arrive in batches and nobody waits on their feedback, so they get one slot; human PRs keep two, in a separate group namespace so a person's PR never queues behind a companion batch.Worst case: 38 cells x 3 slot tokens (
h0,h1,a0) = 114 concurrent Vercel e2e jobs, under the point where CI dominates the ceiling, against the 379 observed at the incident peak.Mechanism
Within a run every cell is a distinct group, so a PR still fans out in full and single-PR latency is unchanged (verified: all 38 lanes on this PR started inside a 2-second window). What collides is the same cell from a different run. The slot token is computed in
ci-scopebecause GitHub expressions have no arithmetic operators; keying onrun_idkeeps a re-run in the slot it already had.Tunable without a code change via
E2E_VERCEL_CONCURRENCY_SLOTSandE2E_VERCEL_CONCURRENCY_SLOTS_AUTOMATED.Tradeoff
Under a burst the eleventh companion PR waits rather than running. That is the point, but it is a latency cost, and head-of-line blocking is real: a lane held by a 30-minute job blocks that lane within its tier. Superseded runs are still cancelled by the workflow-level
cancel-in-progress, which releases their queued jobs, so a re-push frees capacity rather than adding to it.Scope is the four
e2e-vercel-*lanes. The local, postgres, Windows and community lanes generate no traffic against our services.benchmarks.ymlalso deploys and drives load (peak 11 concurrent jobs in the same window) and is not covered here.Verification
E2E Python Conformance, which is red on main's last four runs (19:57 UTC onward) and green at 17:14 — inherited, not from this change.bash --noprofile --norc -eo pipefail) across both tiers and unset / empty /0/4/abcinputs; all exit 0 and fall back conservatively to 1 slot on garbage.Related:
vercel/api#90538raises the HPA ceiling 64 -> 96.🤖 Generated with Claude Code