[world-local] Remove default queue delivery deadlines - #4351
Conversation
A delivery runs inline steps before it responds, so the 30s headers/body deadline redelivered live messages for any step longer than 30s and executed the step a second time. WORKFLOW_LOCAL_HEADERS_TIMEOUT_MS and WORKFLOW_LOCAL_BODY_TIMEOUT_MS remain as opt-in deadlines. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 2e457c6 The changes in this PR will be included in the next version bump. This PR includes changesets to release 18 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 134799ms → this run 156291ms (Δ +21492ms, +16%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): 📜 Previous results (1)b4bb3eeThu, 24 Sep 2026 02:24:31 GMT · run logs
Streams
ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (62 failed)python-node (62 failed):
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ❌ ▲ Vercel Production | 3816 | 62 | 739 | 4617 |
| ✅ 💻 Local Development | 4238 | 0 | 550 | 4788 |
| ✅ 📦 Local Production | 4238 | 0 | 550 | 4788 |
| ✅ 🐘 Local Postgres | 4238 | 0 | 550 | 4788 |
| ✅ 🪟 Windows | 340 | 0 | 2 | 342 |
| ✅ vercel-http-transport | 873 | 0 | 153 | 1026 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 591 | 0 | 93 | 684 |
| Total | 18361 | 62 | 2637 | 21060 |
Details by Category
❌ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 141 | 0 | 30 |
| ✅ astro-quickjs | 141 | 0 | 30 |
| ✅ example-node | 141 | 0 | 30 |
| ✅ example-quickjs | 141 | 0 | 30 |
| ✅ express-node | 141 | 0 | 30 |
| ✅ express-quickjs | 141 | 0 | 30 |
| ✅ fastify-node | 141 | 0 | 30 |
| ✅ fastify-quickjs | 141 | 0 | 30 |
| ✅ hono-node | 141 | 0 | 30 |
| ✅ hono-quickjs | 141 | 0 | 30 |
| ✅ nest-node | 141 | 0 | 30 |
| ✅ nest-quickjs | 141 | 0 | 30 |
| ✅ nextjs-turbopack-node | 168 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 168 | 0 | 3 |
| ✅ nextjs-webpack-node | 168 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 168 | 0 | 3 |
| ✅ nitro-node | 141 | 0 | 30 |
| ✅ nitro-quickjs | 141 | 0 | 30 |
| ✅ nuxt-node | 141 | 0 | 30 |
| ✅ nuxt-quickjs | 141 | 0 | 30 |
| ❌ python-node | 4 | 62 | 105 |
| ✅ sveltekit-node | 160 | 0 | 11 |
| ✅ sveltekit-quickjs | 160 | 0 | 11 |
| ✅ tanstack-start-node | 141 | 0 | 30 |
| ✅ tanstack-start-quickjs | 141 | 0 | 30 |
| ✅ vite-node | 141 | 0 | 30 |
| ✅ vite-quickjs | 141 | 0 | 30 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 170 | 0 | 1 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ hono | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ nitro | 141 | 0 | 30 |
| ✅ vite | 141 | 0 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ vite | 141 | 0 | 30 |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
pranaygp
left a comment
There was a problem hiding this comment.
LGTM. Narrow, correct fix for #3909, and it matches world-postgres (#4114). Both transports treat 0 as "no deadline" (undici natively; nodeHttpFetch via its if (headersTimeoutMs) / if (!bodyTimeoutMs) guards), and close() still aborts in-flight deliveries.
Verified
pnpm --filter @workflow/world-local test: 598 passed, 4 skipped.- End-to-end repro with a real
createQueueand a handler that takes 35s, on both undici andWORKFLOW_NODE_HTTP=1:- New default: 1 delivery, for a stall before headers and for a stall mid-body.
- Opt-in
30000(the old default): 2 deliveries in every case, i.e. #3909 reproduces.
- Node's
http.Server.requestTimeoutdoes not cut off a slow response (requestTimeout=1000with a 4s handler gave 1 delivery), so there's no hidden server-side deadline that would bring the redelivery back.
Follow-ups (not blocking)
- Pre-existing, now more likely: an opt-in value above 2³¹−1 ms breaks every delivery on the
node:httptransport. The new docs tell operators to set the deadline "above the longest inline work you expect", so some will pick a huge value to be safe.envTimeoutMs(queue.ts:82) doesn't clamp, and Node'ssetTimeouttreats an overflowing delay as 1ms (TimeoutOverflowWarning). WithWORKFLOW_NODE_HTTP=1 WORKFLOW_LOCAL_HEADERS_TIMEOUT_MS=3000000000and a 3s handler I got 2no response headers within 3000000000msfailures within about 5s. Every retry dies after 1ms until the 256-attempt safety limit. undici is fine. Fix: clamp toMAX_SAFE_TIMEOUT_MS. world-postgres has the same helper (world-postgres/src/queue.ts:114). - Pre-existing, now more likely:
parseIntmakes plausible typos into tiny deadlines.1e4parses as 1ms,30sas 30ms,30_000as 30ms. That now triggers exactly the duplicate step execution the docs warn about. Suggest strict parsing (Number(raw)plusNumber.isSafeInteger) and a warning when an override is ignored. - Observability: a hung handler is now silent. It holds its semaphore slot forever with no log line. A non-destructive watchdog (a warn or
debugLogafter N minutes namingrunId/stepId/messageId, with no abort) would make stuck runs debuggable. Related: the transport-retry log usesString(err), which on undici printsfetch failedand drops theHeadersTimeoutError/BodyTimeoutErrorcause. Logging the cause and naming the env var would help opt-in users. - Test: consider asserting that the
Agent(or thenodeHttpFetchper-request options) is actually constructed with0/0when no env is set, so the test covers the wiring and not onlygetQueueAgentOptions().
Nits
docs/content/docs/v5/configuration/worlds.mdx: theWORKFLOW_LOCAL_BODY_TIMEOUT_MSentry didn't get the re-execution caution that the headers entry andlocal.mdxhave. Also, neither page now says explicitly that0disables the deadline, which helps someone switching back from an opt-in value.queue.test.ts:561: the first assertion compares against the exported constants, so it can't fail when a constant changes. Asserting{ headersTimeout: 0, bodyTimeout: 0 }directly reads as the spec.queue.ts:79: theDEFAULT_BODY_TIMEOUT_MSdoc comment doesn't say "none".
|
@VaguelySerious here are the regression tests I'd recommend, written and run locally against this branch ( 1. For this PR: pin the wiring, not just
|
Records the options the undici Agent and node:http deliveries are built with, so the test fails if createQueue stops honoring the defaults. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
No backport to This is a genuine correctness fix (duplicate step execution from a 30s queue delivery deadline), but the defect does not exist on To override, re-run the Backport to stable workflow manually via |
Defaults-only alternative to #4328 that fixes #3909.
Problem
Since #3255, world-local gives each queue delivery a 30s undici
headersTimeout/bodyTimeout. A delivery runs inline steps before the handler responds, so any"use step"that takes longer than 30s has its live delivery treated as a transport failure, redelivered, and its body executed a second time while the first is still running. v4 world-local used no custom agent, so it got undici's 300s defaults.Change
DEFAULT_HEADERS_TIMEOUT_MSandDEFAULT_BODY_TIMEOUT_MSgo from30_000to0(no deadline), matching what fix(world-postgres): remove implicit queue HTTP deadlines #4114 did for world-postgres.WORKFLOW_LOCAL_HEADERS_TIMEOUT_MS/WORKFLOW_LOCAL_BODY_TIMEOUT_MSstill work, now as an opt-in. Docs warn that a value below the longest inline step brings the duplication back.This PR leaves out the
runtime.tssingle-flight change from #4328, so it is a lower-risk option for the v5 stable release. #4328 still covers worlds whose transport can redeliver a live message.Tests
pnpm --filter @workflow/world-local test: 602 passed.typecheckclean.0(disabled) configuration. The stalled-handler and stalled-body tests already set their deadlines explicitly, so they now cover the opt-in path.🤖 Generated with Claude Code