Skip to content

[ci] Cap concurrent Vercel E2E lanes repo-wide - #4039

Merged
VaguelySerious merged 1 commit into
mainfrom
peter/ci-e2e-concurrency-cap
Sep 8, 2026
Merged

VaguelySerious merged 1 commit into
mainfrom
peter/ci-e2e-concurrency-cap

Conversation

@VaguelySerious

@VaguelySerious VaguelySerious commented Sep 8, 2026 •

Copy link
Copy Markdown
Member

Problem

The Vercel E2E lanes in tests.yml fan out to 38 jobs per run (27 e2e-vercel-prod, 6 HTTP transport, 4 WS transport, 1 multi-region). Each drives real traffic at api-workflow, and nothing bounded that across runs: the only concurrency control in the file dedupes a branch against itself but lets N open PRs multiply load by N.

That the E2E runner is an api-workflow client is a property of the code path, not an inference. packages/world-vercel/src/utils.ts:

usingProxy = Boolean(projectConfig?.projectId && projectConfig?.teamId)
baseUrl    = usingProxy ? 'https://api.vercel.com/v1/workflow' : `${defaultHost}/api`

CI sets WORKFLOW_VERCEL_PROJECT + WORKFLOW_VERCEL_TEAM + WORKFLOW_VERCEL_AUTH_TOKEN on every Vercel lane, so the test runner takes the proxy. The deployed app under test (OIDC, no projectConfig) goes direct.

Does GitHub support a cap?

Partly, and only since May 2026. There is still no "max N concurrent" setting: a concurrency group admits one running entry. What changed is the queue behind it. queue: max lets up to 100 entries wait instead of the default queue: single, which keeps one pending entry and cancels the rest. Without it, capping a required check would turn a burst of PRs into cancelled jobs and a red E2E Required Check. So the cap is built from many one-at-a-time groups rather than one N-at-a-time group.

queue is documented under workflow-level concurrency and not shown in the jobs.<job_id>.concurrency section, so I probed it: one job, one fixed group, queue: max, three pushes. Two runs sat pending in the group simultaneously (impossible under queue: single), none was cancelled, and they executed strictly FIFO. Job-level queue: max is honored.

Sizing, from INC-7881

Rather than pick a number, I correlated this file's Vercel lanes against cluster_monitor_production.hpa_desired_replicas{hpa_name:api-workflow,env:sfo1} for 2026-09-08 19:00–22:30 UTC, per minute:

concurrent Vercel e2e jobs minutes at the 64-replica ceiling
0 (CI idle) 41%
1–50 41%
51–150 62%
151–250 94%
251+ 100%

Pearson r = 0.43.

Two things follow, and the first is a caveat on this PR. The HPA reaches its ceiling 41% of the time with CI completely idle, and the earlier 33-minute saturation at 19:16–19:49 (which also crossed the monitor's 30-minute window) began with zero e2e jobs running. CI is an aggravating load, not the whole story, and this change alone will not keep api-workflow off its ceiling. Second: CI's own contribution stops being marginal somewhere above 150 concurrent jobs, which is what sizes the tiers.

What the incident peak was made of

At 20:59, the minute the HPA pinned, all 379 concurrent Vercel e2e jobs came from ten test/server-pr-* branches — the companion PRs a workflow-server PR opens here — each running the full 38-cell fan-out. Not one was a human PR.

So the cap is tiered. Machine-opened branches (test/server-pr-*, changeset-release/*) arrive in batches and nobody waits on their feedback, so they get one slot; human PRs keep two, in a separate group namespace so a person's PR never queues behind a companion batch.

Worst case: 38 cells x 3 slot tokens (h0, h1, a0) = 114 concurrent Vercel e2e jobs, under the point where CI dominates the ceiling, against the 379 observed at the incident peak.

Mechanism

concurrency:
  group: e2e-vercel-prod-${{ matrix.app.name }}-${{ matrix.vm }}-${{ needs.ci-scope.outputs.e2e-vercel-slot }}
  queue: max

Within a run every cell is a distinct group, so a PR still fans out in full and single-PR latency is unchanged (verified: all 38 lanes on this PR started inside a 2-second window). What collides is the same cell from a different run. The slot token is computed in ci-scope because GitHub expressions have no arithmetic operators; keying on run_id keeps a re-run in the slot it already had.

Tunable without a code change via E2E_VERCEL_CONCURRENCY_SLOTS and E2E_VERCEL_CONCURRENCY_SLOTS_AUTOMATED.

Tradeoff

Under a burst the eleventh companion PR waits rather than running. That is the point, but it is a latency cost, and head-of-line blocking is real: a lane held by a 30-minute job blocks that lane within its tier. Superseded runs are still cancelled by the workflow-level cancel-in-progress, which releases their queued jobs, so a re-push frees capacity rather than adding to it.

Scope is the four e2e-vercel-* lanes. The local, postgres, Windows and community lanes generate no traffic against our services. benchmarks.yml also deploys and drives load (peak 11 concurrent jobs in the same window) and is not covered here.

Verification

  • First push: all 38 Vercel lanes green. The only failure was E2E Python Conformance, which is red on main's last four runs (19:57 UTC onward) and green at 17:14 — inherited, not from this change.
  • Slot script exercised under GitHub's exact shell flags (bash --noprofile --norc -eo pipefail) across both tiers and unset / empty / 0 / 4 / abc inputs; all exit 0 and fall back conservatively to 1 slot on garbage.

Related: vercel/api#90538 raises the HPA ceiling 64 -> 96.

🤖 Generated with Claude Code

The Vercel E2E lanes in tests.yml fan out to 38 jobs per run (27 in
e2e-vercel-prod, 6 HTTP transport, 4 WS transport, 1 multi-region), and
every one of them drives real traffic at world-vercel and at a shared
Vercel project. Nothing bounded that across runs: the only concurrency
control in the file is the workflow-level group, which dedupes a branch
against itself but lets N open PRs multiply the load by N.

Give each Vercel lane a job-level concurrency group keyed on (lane,
matrix cell, slot). Within one run every cell is a distinct group, so a
single PR still fans out in full; the same cell from a different run
queues instead. The slot count, computed in ci-scope because GitHub
expressions have no arithmetic, sets how many runs' worth may be in
flight repo-wide and is tunable via the E2E_VERCEL_CONCURRENCY_SLOTS
repository variable (default 2).

`queue: max` is required rather than incidental. The default
`queue: single` holds one pending entry per group and cancels the rest,
which would convert a burst of PRs into cancelled jobs and a red E2E
Required Check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
example-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workflow-docs Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC
workflow-web Ready Ready Preview, v0 Sep 8, 2026 9:41pm UTC

@changeset-bot

changeset-bot Bot commented Sep 8, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 3e49a51

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • addTenWorkflow (express)
  • addTenWorkflow (vite)
  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • experimental_retention: 0 purges the run payloads once the run finishes (express)
  • hookAdoptOwnerResultWorkflow - duplicate adopts the owner result via conflict.returnValue (nest)
  • promiseAllWorkflow (fastify)
  • promiseRaceStressTestWorkflow (fastify)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • cold-start-warmup · suite warmup (tanstack-start) · at 21:41:01Z · abandoned wrun_01M21FED653KKZ9TRNGRTET876
  • run-pickup-stall · hookSignalOwnerWorkflow - duplicate forwards its payload to the owner via resumeHook (nuxt) · at 21:42:26Z · abandoned wrun_01M21FHBV7KF8R6YA32A46RQMZ

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3662 0 685 4347
✅ 💻 Local Development 3922 0 586 4508
✅ 📦 Local Production 3922 0 586 4508
✅ 🐘 Local Postgres 3922 0 586 4508
✅ 🪟 Windows 320 0 2 322
✅ 🌐 Cross-language Conformance 68 0 74 142
✅ vercel-http-transport 823 0 143 966
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 557 0 87 644
Total 17223 0 2749 19972
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 133 0 28
✅ astro-quickjs 133 0 28
✅ example-node 133 0 28
✅ example-quickjs 133 0 28
✅ express-node 133 0 28
✅ express-quickjs 133 0 28
✅ fastify-node 133 0 28
✅ fastify-quickjs 133 0 28
✅ hono-node 133 0 28
✅ hono-quickjs 133 0 28
✅ nest-node 133 0 28
✅ nest-quickjs 133 0 28
✅ nextjs-turbopack-node 158 0 3
✅ nextjs-turbopack-quickjs 158 0 3
✅ nextjs-webpack-node 158 0 3
✅ nextjs-webpack-quickjs 158 0 3
✅ nitro-node 133 0 28
✅ nitro-quickjs 133 0 28
✅ nuxt-node 133 0 28
✅ nuxt-quickjs 133 0 28
✅ python-node 66 0 95
✅ sveltekit-node 152 0 9
✅ sveltekit-quickjs 152 0 9
✅ tanstack-start-node 133 0 28
✅ tanstack-start-quickjs 133 0 28
✅ vite-node 133 0 28
✅ vite-quickjs 133 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 141 0 20
✅ nextjs-turbopack-canary-quickjs 141 0 20
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 141 0 20
✅ nextjs-webpack-canary-quickjs 141 0 20
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 141 0 20
✅ nextjs-turbopack-canary-quickjs 141 0 20
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 141 0 20
✅ nextjs-webpack-canary-quickjs 141 0 20
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 141 0 20
✅ nextjs-turbopack-canary-quickjs 141 0 20
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 141 0 20
✅ nextjs-webpack-canary-quickjs 141 0 20
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 160 0 1
✅ nextjs-turbopack-quickjs 160 0 1

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 74

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 133 0 28
✅ express 133 0 28
✅ hono 133 0 28
✅ nextjs-turbopack 158 0 3
✅ nitro 133 0 28
✅ vite 133 0 28

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 133 0 28
✅ express 133 0 28
✅ nextjs-turbopack 158 0 3
✅ vite 133 0 28

📋 View full workflow run

@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 3e49a51 · Tue, 08 Sep 2026 22:03:40 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 176 (-74%) 💚 1236 🔴 (+23%) 🔻 1263 🔴 (+25%) 🔻 1405 🔴 (+7.9%) 30
TTFS stream 188 (+1.6%) 1291 🔴 (+29%) 🔻 1327 🔴 (+29%) 🔻 1600 🔴 (+36%) 🔻 30
TTFS hook + stream 1453 (+248%) 🔻 1600 🔴 (+28%) 🔻 1720 🔴 (+33%) 🔻 1950 🔴 (+18%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 674 (+44%) 🔻 2050 (+129%) 🔻 2131 (+90%) 🔻 2256 (+62%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 1956 (+10%) 4196 (+13%) 4625 (+5.9%) 6405 (-12%) 10
STSO 1020 steps (inline) 107 (+19%) 🔻 134 (+12%) 161 (+17%) 🔻 247 (+36%) 🔻 1019
WO 1020 steps 136282 (+13%) 136282 (+13%) 136282 (+13%) 136282 (+13%) 1
CRTT first chunk (pooled) 59 (-4.8%) 91 (-7.1%) 153 (-18%) 💚 173 (-35%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 83.5 (+11%) 356 (+151%) 739 (+281%) 3565 (+927%) 208 (+47%) 10
size sweep (100/s, 160B-12KB) 79 (-14%) 172 (+4%) 359 (+9%) 744 (-38%) 143 (-32%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 82 (-39%) 131 (-2%) 209 (+11%) 769 (+148%) 220 (-7%) 3
replay eve-gpt-5.6-sol-2000t (1x) 94.5 (+28%) 200 (+38%) 2263 (+906%) 3600 (+550%) 2233 (+447%) 2
replay eve-gpt-5.6-sol-2000t (2x) 81 (+3%) 236 (-8%) 347 (-33%) 751 (-42%) 586 (+15%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 120942ms → this run 135981ms (Δ +15039ms, +12%)

 50-100 ms  ┃                         main   2  this   0    -2
100-150 ms  ████████████████████┃███  main 968  this 863  -105
150-200 ms  █░┃                       main  44  this 126   +82
200-250 ms  ┃                         main   2  this  21   +19
250-300 ms  ┃                         main   2  this   4    +2
300-350 ms  ┃                         main   0  this   3    +3
350-400 ms  ┃                         main   0  this   1    +1
500-550 ms  ┃                         main   1  this   0    -1
550-600 ms  ┃                         main   0  this   1    +1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+              avg         p50           p90           p99     n
control  ······▄█▃▁·▁·  299.4 (+155%)  125 (+19%)   739 (+281%)  3565 (+927%)  3000
sweep    ······▃█▂▁···   144.7 (-10%)   119 (+6%)     359 (+9%)    744 (-38%)  3000
gw 1x    ·····▁██▁▁···    116.4 (+7%)   101 (+6%)    209 (+11%)   769 (+148%)  5295
eve 1x   ·····▁█▇▂▁▁▁·  298.8 (+139%)    94 (-4%)  2263 (+906%)  3600 (+550%)  5186
eve 2x   ·····▁▃█▄▁···   178.6 (-17%)   152 (+3%)    347 (-33%)    751 (-42%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▂▂▁▁▂▂█▇▆▅  156–554ms
sweep    █▄▁▁▅▆▄█▅█  121–163ms
gw 1x    ▃█▂▃▃▁▁▁▁▃  96–183ms
eve 1x   ▁▁▁▁▁▁▄█▂▁  92–1345ms
eve 2x   ▂▂▅▆▁▁▂█▇▁  134–259ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▂█▄▅▆▁▃  143–147ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▄▂▂▂▂█▃▁▁▂  34–122ms
sweep    ▂▁▃▃▄▄█▂▅▄  47–70ms
gw 1x    ▄█▁▄▃▁▅▂▁▄  28–42ms
eve 1x   ▂▁▁▂▂▂█▄▃▂  20–49ms
eve 2x   ▄▆▃█▃▃▂▁▅▅  22–34ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 201.6 KiB (±0) 41.7 KiB (±0) 1.78 MiB (±0)
nextjs-turbopack 207.0 KiB (±0) 439 B (±0) 770.5 KiB (±0)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

3e49a51 · run

@VaguelySerious
VaguelySerious merged commit 9a9af61 into main Sep 8, 2026
181 of 182 checks passed
@VaguelySerious
VaguelySerious deleted the peter/ci-e2e-concurrency-cap branch September 8, 2026 21:47
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

No backport to stable for 9a9af61 (AI decision).

This adds a new CI capacity-management capability (per-cell/slot job concurrency groups plus a tunable E2E_VERCEL_CONCURRENCY_SLOTS repository variable) to throttle Vercel E2E fan-out; it is an infrastructure enhancement and throughput/latency trade, not a fix for CI that is broken on stable. stable remains buildable, testable, and releasable without it, and it ships no user-facing package change (the changeset is empty). If Vercel backend load from stable runs becomes a real problem, this can be forced through via workflow_dispatch.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

9a9af618f7e5e9d5a27fb8682860c736c7980dae

pranaygp added a commit that referenced this pull request Sep 9, 2026
…c-workflow-source

* origin/main: (22 commits)
  feat(streams): add writer session seam (#3832)
  docs: document WORKFLOW_NODE_HTTP in the v4 World docs (#4050)
  [world-vercel] Honor WORKFLOW_NODE_HTTP on the queue transport (#4044)
  feat(streams): add WebSocket capability gate (#3764)
  fix(world-local): retry JSON reads on Windows (#4051)
  [core] Add the wake-loop scenario to the event log race repro (#4017)
  [ci] Cap concurrent Vercel E2E action repo-wide (10 by default) (#4039)
  Add `Run#getWritable()` for appending to another run's stream (#3972)
  feat(streams): add WebSocket v1 client protocol contract (#3763)
  [swc-playground] Update to Next.js v16.3.4 (#4018)
  [core] Add a retention option to start() (#3787)
  fix(swc-plugin): register class expressions via an IIFE instead of by name (#3971)
  [docs] Fix prose typos across v4/v5 docs and the SWC plugin README (#3948)
  Validate pending changesets in CI so a bad one fails the PR, not the Release job (#3964)
  Version Packages (beta) (#3919)
  Classify Workflow stream failures (#3850)
  Drop the ignored @workflow/world-sim package from the hook_conflict delta changeset (#3963)
  Add attribute inspection to the CLI (#3950)
  [core] Settle a hook's awaiter in-process instead of re-invoking, on creation and on conflict (#3938)
  Use the storage APIs for run detail views (#3944)
  ...

This branch was successfully deployed

18 active deployments
Preview – workflow-swc-playground — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workflow-docs — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-vite-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – example-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-astro-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workflow-tarballs — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-hono-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-express-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workflow-web — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Preview – workbench-python-workflow — 3e49a513 Deployed Sep 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants