fix(world-vercel): keep large event bodies off busy HTTP/2 connections - #4458
Conversation
Event POSTs with bodies over 1 MiB (run_started carrying a multi-MiB run input, large event batches) could not be re-buffered by the H2 multiplexing interceptor, so undici held them until the connection had zero in-flight streams and stopped its queue behind them. On busy instances they hung until the replay budget or maxDuration ran out. - Send those bodies over an HTTP/1.1 agent instead. - Upgrade undici to 7.30.0, which fixes the phantom running slot a failed HTTP/2 stream left on its connection (nodejs/undici#5410, #5569) and honors headersTimeout on HTTP/2. - Bound events requests with 60s header/body deadlines instead of undici's 300s defaults, which outlasted the 240s replay budget. - Retry event-log reads (GET only) whose stream the peer reset with ERR_HTTP2_STREAM_ERROR. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 6e550fe The changes in this PR will be included in the next version bump. This PR includes changesets to release 19 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 176751ms → this run 152282ms (Δ -24469ms, -14%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (26 failed)astro-node (1 failed):
astro-quickjs (1 failed):
example-node (1 failed):
example-quickjs (1 failed):
express-node (1 failed):
express-quickjs (1 failed):
fastify-node (1 failed):
fastify-quickjs (1 failed):
hono-node (1 failed):
hono-quickjs (1 failed):
nest-node (1 failed):
nest-quickjs (1 failed):
nextjs-turbopack-node (1 failed):
nextjs-turbopack-quickjs (1 failed):
nextjs-webpack-node (1 failed):
nextjs-webpack-quickjs (1 failed):
nitro-node (1 failed):
nitro-quickjs (1 failed):
nuxt-node (1 failed):
nuxt-quickjs (1 failed):
sveltekit-node (1 failed):
sveltekit-quickjs (1 failed):
tanstack-start-node (1 failed):
tanstack-start-quickjs (1 failed):
vite-node (1 failed):
vite-quickjs (1 failed):
💻 Local Development (1 failed)nextjs-webpack-stable-node (1 failed):
vercel-http-transport (6 failed)example (1 failed):
express (1 failed):
hono (1 failed):
nextjs-turbopack (1 failed):
nitro (1 failed):
vite (1 failed):
vercel-ws-transport (4 failed)example (1 failed):
express (1 failed):
nextjs-turbopack (1 failed):
vite (1 failed):
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ❌ ▲ Vercel Production | 3852 | 26 | 739 | 4617 |
| ❌ 💻 Local Development | 4237 | 1 | 550 | 4788 |
| ✅ 📦 Local Production | 4238 | 0 | 550 | 4788 |
| ✅ 🐘 Local Postgres | 4238 | 0 | 550 | 4788 |
| ✅ 🪟 Windows | 340 | 0 | 2 | 342 |
| ✅ 🌐 Cross-language Conformance | 68 | 0 | 84 | 152 |
| ❌ vercel-http-transport | 867 | 6 | 153 | 1026 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ❌ vercel-ws-transport | 587 | 4 | 93 | 684 |
| Total | 18454 | 37 | 2721 | 21212 |
Details by Category
❌ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ❌ astro-node | 140 | 1 | 30 |
| ❌ astro-quickjs | 140 | 1 | 30 |
| ❌ example-node | 140 | 1 | 30 |
| ❌ example-quickjs | 140 | 1 | 30 |
| ❌ express-node | 140 | 1 | 30 |
| ❌ express-quickjs | 140 | 1 | 30 |
| ❌ fastify-node | 140 | 1 | 30 |
| ❌ fastify-quickjs | 140 | 1 | 30 |
| ❌ hono-node | 140 | 1 | 30 |
| ❌ hono-quickjs | 140 | 1 | 30 |
| ❌ nest-node | 140 | 1 | 30 |
| ❌ nest-quickjs | 140 | 1 | 30 |
| ❌ nextjs-turbopack-node | 167 | 1 | 3 |
| ❌ nextjs-turbopack-quickjs | 167 | 1 | 3 |
| ❌ nextjs-webpack-node | 167 | 1 | 3 |
| ❌ nextjs-webpack-quickjs | 167 | 1 | 3 |
| ❌ nitro-node | 140 | 1 | 30 |
| ❌ nitro-quickjs | 140 | 1 | 30 |
| ❌ nuxt-node | 140 | 1 | 30 |
| ❌ nuxt-quickjs | 140 | 1 | 30 |
| ✅ python-node | 66 | 0 | 105 |
| ❌ sveltekit-node | 159 | 1 | 11 |
| ❌ sveltekit-quickjs | 159 | 1 | 11 |
| ❌ tanstack-start-node | 140 | 1 | 30 |
| ❌ tanstack-start-quickjs | 140 | 1 | 30 |
| ❌ vite-node | 140 | 1 | 30 |
| ❌ vite-quickjs | 140 | 1 | 30 |
❌ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ❌ nextjs-webpack-stable-node | 169 | 1 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 170 | 0 | 1 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 68 | 0 | 84 |
❌ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ❌ example | 140 | 1 | 30 |
| ❌ express | 140 | 1 | 30 |
| ❌ hono | 140 | 1 | 30 |
| ❌ nextjs-turbopack | 167 | 1 | 3 |
| ❌ nitro | 140 | 1 | 30 |
| ❌ vite | 140 | 1 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
❌ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ❌ example | 140 | 1 | 30 |
| ❌ express | 140 | 1 | 30 |
| ❌ nextjs-turbopack | 167 | 1 | 3 |
| ❌ vite | 140 | 1 | 30 |
About these numbersSizes are gzip; parentheses show the change against
|
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
#4458) Event POSTs with bodies over 1 MiB (run_started carrying a multi-MiB run input, large event batches) could not be re-buffered by the H2 multiplexing interceptor, so undici held them until the connection had zero in-flight streams and stopped its queue behind them. On busy instances they hung until the replay budget or maxDuration ran out. - Send those bodies over an HTTP/1.1 agent instead. - Upgrade undici to 7.30.0, which fixes the phantom running slot a failed HTTP/2 stream left on its connection (nodejs/undici#5410, #5569) and honors headersTimeout on HTTP/2. - Bound events requests with 60s header/body deadlines instead of undici's 300s defaults, which outlasted the 240s replay budget. - Retry event-log reads (GET only) whose stream the peer reset with ERR_HTTP2_STREAM_ERROR. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
|
Backport PR opened against |
#4458) (#4461) Event POSTs with bodies over 1 MiB (run_started carrying a multi-MiB run input, large event batches) could not be re-buffered by the H2 multiplexing interceptor, so undici held them until the connection had zero in-flight streams and stopped its queue behind them. On busy instances they hung until the replay budget or maxDuration ran out. - Send those bodies over an HTTP/1.1 agent instead. - Upgrade undici to 7.30.0, which fixes the phantom running slot a failed HTTP/2 stream left on its connection (nodejs/undici#5410, #5569) and honors headersTimeout on HTTP/2. - Bound events requests with 60s header/body deadlines instead of undici's 300s defaults, which outlasted the 240s replay budget. - Retry event-log reads (GET only) whose stream the peer reset with ERR_HTTP2_STREAM_ERROR. Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Problem
Under high flow volume (the ai-gateway content-capture workflow on 2026-09-28), some deliveries stalled for 5–13 minutes and then retried:
run_startedcarrying a multi-MiB run input, or a largeevents/batch. In the workflow-server logs these requests never reached the proxy. The same POST was re-sent exactly 300 s later, or the invocation was killed atmaxDuration(800 s) without sending anything. Traced runs:wrun_41M3MJAJPY50SGJG99W67P7287(23 MBrun_started),wrun_41M3MH64AE20XC2W8GMV1MRMQM(2.27 MB batch),wrun_41M3MAS0WV20VDYHGXCPRHCTFE(1.45 MBrun_started, 28 minutes end to end).GET listEventsfailing withERR_HTTP2_STREAM_ERROR(NGHTTP2_ENHANCE_YOUR_CALM/NGHTTP2_INTERNAL_ERRORfrom the edge). workflow-server logged these requests as 200. They were never retried in-process, so each one cost a full redelivery.Root cause
Streamed bodies starve on a busy HTTP/2 connection.
client-h2.jsbusy()holds a streamed body until the connection has zero running streams.client.js_resumethen stops dispatching the whole queue behind it.headersTimeoutonly arms once a request reaches a socket.undici < 7.30.0 leaks a "running" slot on every failed HTTP/2 stream (fixed by the backport of fix: preserve h2 queue on out-of-order completion nodejs/undici#5410 and #5569). After a single peer
RST_STREAM, the connection reportsrunning: 1forever. From then on, no streamed body is ever dispatched on it again. That links the edge resets to the stuck large POSTs. Minimal repro, the same script against each version:runningafter one reset7.30.0 also makes
headersTimeoutapply on HTTP/2; before, onlybodyTimeoutwas used for both.undici's 300 s defaults outlast the 240 s replay budget. A silent stream ran past the replay timeout before it failed.
The events RetryAgent doesn't retry
ERR_HTTP2_STREAM_ERROR. It isn't in undici's defaulterrorCodes.Changes
createH2MultiplexInterceptor(fallback)sends bodies it won't re-buffer (over 1 MiB, or unknown length) to an HTTP/1.1 agent owned by the events dispatcher.h2MultiplexInterceptoris still exported, unchanged, with no fallback.EVENTS_REQUEST_TIMEOUT_MS = 60_000. Used asheadersTimeout/bodyTimeouton both events agents, matchingREQUEST_TIMEOUT_MSon the v3 path. It bounds silence, not total duration, so long streamed pages are unaffected.EVENTS_RETRY_AGENT_OPTIONS. undici's defaulterrorCodesplusERR_HTTP2_STREAM_ERROR.methodsstays at undici's default (no POST), so event writes are still never replayed.Tests
New
describe('events dispatcher on a busy HTTP/2 connection'), run against a loopback h2 origin withallowHTTP1:createEventsDispatchersends the same POST over HTTP/1.1 and it completes while the stream is still held.ENHANCE_YOUR_CALMis retried and succeeds.running: 0. This fails on 7.29.0.pnpm vitest runin world-vercel: 766 passed. world-local: 604 passed.tsc --noEmitis clean.Rollout
ai-gateway (
workflow@5.0.0-beta.53) will take this in a version bump instead of settingWORKFLOW_H2_MULTIPLEX=0.Related: vercel/vqs-server#677 fixes the other half of the stall, where any retry directive waited out the 300 s lease.
🤖 Generated with Claude Code