Repository navigation
perf(world-vercel): batch a fan-out's step-execution queue publishes - #3838
Conversation
A `Promise.all` fan-out dispatched one queue message per branch. Those publishes ride the shared default undici agent (8 connections, HTTP/1.1, `pipelining: 1` — see `getQueueDispatcher`), and `handleSuspension` is awaited in full before the first inline step body runs, so an N-branch fan-out paid ~N/8 serialized round trips straight onto time-to-first-step. The `step_created` writes were already batched and HTTP/2-multiplexed; the publishes were the remaining per-branch round trip. Adds an optional `Queue.queueBatch`, implemented on `@vercel/queue`'s `experimental_sendBatch` (0.5.1), and uses it for the batched fan-out fold's publishes. Each commit chunk now publishes in one request instead of up to 32. `queueBatch` reports per-entry outcomes rather than throwing, because a batch can partially fail. `queueMessages` in core keeps the previous all-or-nothing behavior for this call site: it rejects if any entry failed, so the delivery is redelivered and republishes the set, deduped by the per-step `idempotencyKey` the caller already passed. Worlds without `queueBatch` fall back to concurrent single sends. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 367e5e0 The changes in this PR will be included in the next version bump. This PR includes changesets to release 20 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ✅ ▲ Vercel Production | 3662 | 0 | 685 | 4347 |
| ✅ 💻 Local Development | 3922 | 0 | 586 | 4508 |
| ✅ 📦 Local Production | 3922 | 0 | 586 | 4508 |
| ✅ 🐘 Local Postgres | 3922 | 0 | 586 | 4508 |
| ✅ 🪟 Windows | 320 | 0 | 2 | 322 |
| ✅ 🌐 Cross-language Conformance | 68 | 0 | 74 | 142 |
| ✅ vercel-http-transport | 823 | 0 | 143 | 966 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 557 | 0 | 87 | 644 |
| Total | 17223 | 0 | 2749 | 19972 |
Details by Category
✅ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 133 | 0 | 28 |
| ✅ astro-quickjs | 133 | 0 | 28 |
| ✅ example-node | 133 | 0 | 28 |
| ✅ example-quickjs | 133 | 0 | 28 |
| ✅ express-node | 133 | 0 | 28 |
| ✅ express-quickjs | 133 | 0 | 28 |
| ✅ fastify-node | 133 | 0 | 28 |
| ✅ fastify-quickjs | 133 | 0 | 28 |
| ✅ hono-node | 133 | 0 | 28 |
| ✅ hono-quickjs | 133 | 0 | 28 |
| ✅ nest-node | 133 | 0 | 28 |
| ✅ nest-quickjs | 133 | 0 | 28 |
| ✅ nextjs-turbopack-node | 158 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 158 | 0 | 3 |
| ✅ nextjs-webpack-node | 158 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 158 | 0 | 3 |
| ✅ nitro-node | 133 | 0 | 28 |
| ✅ nitro-quickjs | 133 | 0 | 28 |
| ✅ nuxt-node | 133 | 0 | 28 |
| ✅ nuxt-quickjs | 133 | 0 | 28 |
| ✅ python-node | 66 | 0 | 95 |
| ✅ sveltekit-node | 152 | 0 | 9 |
| ✅ sveltekit-quickjs | 152 | 0 | 9 |
| ✅ tanstack-start-node | 133 | 0 | 28 |
| ✅ tanstack-start-quickjs | 133 | 0 | 28 |
| ✅ vite-node | 133 | 0 | 28 |
| ✅ vite-quickjs | 133 | 0 | 28 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 160 | 0 | 1 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 68 | 0 | 74 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 28 |
| ✅ express | 133 | 0 | 28 |
| ✅ hono | 133 | 0 | 28 |
| ✅ nextjs-turbopack | 158 | 0 | 3 |
| ✅ nitro | 133 | 0 | 28 |
| ✅ vite | 133 | 0 | 28 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 28 |
| ✅ express | 133 | 0 | 28 |
| ✅ nextjs-turbopack | 158 | 0 | 3 |
| ✅ vite | 133 | 0 | 28 |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 163045ms → this run 177564ms (Δ +14519ms, +9%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): 📜 Previous results (2)480b34fFri, 11 Sep 2026 17:11:54 GMT · run logs
Streams
68b61acThu, 27 Aug 2026 00:26:58 GMT · run logs
Streams
ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
karthikscale3
left a comment
There was a problem hiding this comment.
AI Review: Reviewed the change and load-tested the batched path at 64 branches. Design is sound and the error semantics are the right ones — per-entry outcomes rather than throws, rejection reserved for request-level failures, idempotency keys documented as a requirement rather than a suggestion. One thing I'd fix before merge, the rest are notes.
Verified independently: @vercel/queue 0.5.0 → 0.5.1 is purely additive (no changes to the single-send path); createQueue's result is spread into the World so queueBatch survives composition; the grouping key covers every routing dimension prepareSend derives; and the prepareSend/clientFor refactor is behavior-preserving for queue().
Wire-level, against a fake VQS over real undici (real createQueue, real SDK, real multipart — nothing mocked below createQueue): 64 branches go out in one request, keys and order intact on the wire; per-entry 429/400 are reported without rejecting; request-level 500 rejects; 140 messages split [100, 40] with global order preserved; and a full republish after a partial failure dedupes 61 and delivers 64 unique. The proxy path handles the /batch suffix and the chunked multipart body fine (client side — this says nothing about api.vercel.com's own allowlist, worth confirming separately if a proxy-mode caller ever adopts queueBatch).
Wedge hunt at the runtime level, driving the real fold with a redelivery loop and a VQS dedupe model, every scenario also run against a world without queueBatch as a control:
| batched | control (pre-PR path) | |
|---|---|---|
| happy 64 | 63 published in 2 requests [32,31], 1 delivery |
63 requests |
| transient entry failures | converges in 2, 0 duplicate steps | same |
| permanent entry failure | never converges, 96 requests | never converges, 3024 requests |
| flaky request-level 503 | converges in 2 | same |
| 3x total outage | converges in 4, events committed once | same |
No wedge is introduced by batching — every failure mode I could reach has an identical outcome on the control, and in the bad case batching is strictly better (96 requests vs 3024 to reach the same dead end). Zero duplicate step deliveries anywhere; the idempotency keys do their job. deferredBatchWork correctly surfaces a failed batched publish rather than swallowing a chunk.
One heads-up for whoever owns the dashboards: fan-out queue.publish span volume drops ~32x and rpc.method becomes publishBatch. messaging.batch.message_count recovers the count, which is the right call, but any monitor counting those spans will see a step change.
| }, | ||
| async () => { | ||
| const results = await batch(queueName, messages); | ||
| const failures = results.filter((result) => result.error !== undefined); |
There was a problem hiding this comment.
AI Review: queueMessages only inspects error, never results.length. A World whose queueBatch returns a short array reports success here: handleSuspension resolves, the delivery is acked, and those steps are never dispatched — the run wedges with no error anywhere in the system.
world-vercel guards this internally (toBatchResult(undefined)) and the SDK length-checks too, so it's unreachable through the world in this PR. But queueBatch is documented in building-a-world.mdx for third-party worlds, which makes it an unenforced contract at exactly the boundary that publishes it — and unlike a rejection, this failure is silent.
I reproduced it with a 64-branch fan-out against a world returning half the results: 32 of 63 steps silently lost, handleSuspension resolved without error. Adding the length check flips it to a clean rejection that converges on redelivery, and all 143 tests in helpers.test.ts + suspension-handler.test.ts still pass:
const results = await batch(queueName, messages);
if (results.length !== messages.length) {
throw Object.assign(
new Error(
`Queue batch for ${queueName} returned ${results.length} result(s) ` +
`for ${messages.length} message(s)`
),
{ retryable: true }
);
}| ); | ||
| // Surfaced so a caller (and the delivery-level retry above it) can tell | ||
| // a transient partial batch from a permanent rejection. | ||
| Object.assign(error, { retryable }); |
There was a problem hiding this comment.
AI Review: retryable is never read anywhere in packages/core/src/runtime, so the comment above it overstates what happens: queueMessages throws identically for retryable and non-retryable, and the delivery retries either way.
That has a measurable cost. With one permanently-rejected entry in a 64-branch fan-out, the delivery burns all 48 redeliveries and ~2,900 redundant republishes before giving up. The unbatched path does the same, so this isn't a regression — but the batched path is the one that now has the information and discards it, which makes a fast-fail on retryable: false cheap to add later.
| >(); | ||
| for (const [index, entry] of messages.entries()) { | ||
| const prepared = prepareSend(queueName, entry.message, entry.opts); | ||
| const key = `${prepared.region} ${prepared.deploymentId} ${prepared.useCbor} ${prepared.topic}`; |
There was a problem hiding this comment.
AI Review: This grouping is a no-op under WORKFLOW_SEQUENTIAL_REPLAYS=1. stepDispatch.queueName is the flow topic, which matches FLOW_TOPIC_PATTERN, so getPhysicalQueueName returns ${queue}_${runId}_${stepId} — a distinct topic per message, hence a distinct group per message, hence N requests of one message each.
Off by default so there's no production impact, but two consequences worth a line of comment: the win silently disappears for opted-in users, and each send now goes through the multipart batch endpoint rather than send(), which loses the SDK's 502 consumer_discovery_failed → ConsumerDiscoveryError mapping (sendBatch only maps 503). A client.send fast path for single-entry groups would restore both.
| process.env.VERCEL_DEPLOYMENT_ID = 'dpl_batch'; | ||
| }); | ||
| afterEach(() => { | ||
| process.env.VERCEL_DEPLOYMENT_ID = undefined; |
There was a problem hiding this comment.
AI Review: Nit: process.env.VERCEL_DEPLOYMENT_ID = undefined sets the string "undefined"; the rest of the file uses delete. Harmless as the last describe in the file, just inconsistent.
`experimental_sendBatch` injects the active trace context into the multipart REQUEST headers, and the per-part headers it builds never see it. VQS stores headers per message and re-emits a stored `traceparent` at delivery as `x-vercel-queue-traceparent`, which is what lets a consumer attach a span link back to its producer, so a batched message arrived with no producer context and its `vqs.process` span got no link. `send()` is unaffected: for a single message the request headers ARE that message's headers. At 64 branches that was 63 of 64 step dispatches losing the transport-level producer link. The run's own step tracing was never affected: that carrier travels in the message payload (`WorkflowInvokePayload.traceCarrier`), which is what the consumer builds its trace context from, not a header. Injects the active context into each entry's headers in `queueBatch` — last, so it wins over caller-supplied `opts.headers` exactly as the SDK's own injection does — and honors VERCEL_QUEUE_TRACE_PROPAGATION so that kill switch still covers both paths. `getTraceContextHeaders()` is factored out of `injectTraceContextIntoHeaders` so the two share one source. Verified on the wire against a stub VQS speaking the real batch endpoint: `traceparent` carrying the producer's traceId/spanId lands on all 64 multipart parts through the real SDK, with the per-message idempotency keys still alongside it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
No backport to This is a performance optimization that adds new API surface: an optional To override, re-run the Backport to stable workflow manually via |
Why
Investigating a durabench report that TTLS (time-to-last-step) grows sharply past 64 parallel branches, I measured where a fan-out's time actually goes. Dispatch is not the problem — all 256
step_createdevents land within one second, because that path is already batched (createBatch, 32-event chunks) and HTTP/2-multiplexed.The queue publishes were not. The batched fan-out fold published one message per step:
Those sends deliberately ride the shared default undici agent —
connections: 8,pipelining: 1,allowH2: false(seegetQueueDispatcher) — so an N-branch fan-out costs roughly N/8 serialized HTTP/1.1 round trips. AndhandleSuspensionis awaited in full before any inline step body runs (runtime.ts:3212→4126), so that cost lands directly on time-to-first-step rather than overlapping with useful work.Measured isolated TTFS medians on the current beta, one run at a time on an idle deployment: 64 branches 505 ms, 128 branches 633 ms, 256 branches 785 ms.
What
@vercel/queue0.5.0 → 0.5.1, which addsexperimental_sendBatch(up to 100 messages per request, per-entry results).Queue.queueBatch(optional) on@workflow/world. Worlds that don't implement it are unaffected.queueBatchin@workflow/world-vercel, onexperimental_sendBatch. It groups messages by the routing dimensions one VQS request cannot span (region / deploymentId / transport / physical topic) and splits at the 100-message cap; neither split is observable in the returned order. For the case this exists for — one run's fan-out to one logical queue — that is a single group.queueMessagesin@workflow/core, used by the fan-out fold'spublishChunkSteps. Each commit chunk now publishes in one request instead of up to 32.Error semantics
queueBatchreports per-entry outcomes rather than throwing, because a batch can partially fail; it rejects only for a request-level failure where no entry outcome is known.queueMessageskeeps this call site's existing all-or-nothing behavior: it rejects if any entry failed, so the delivery is redelivered and republishes the whole set. That is safe precisely because the caller already passes a per-stepidempotencyKey(stepDispatchIdempotencyKey) — without it, republishing would redeliver the steps that already succeeded. The interface docs make that a requirement rather than a suggestion.messageId: nullwith noerroris an acceptance (deferred, no ID yet), matchingqueue(), so success iserror === undefinedrather than a non-null id.Worlds without
queueBatchfall back to concurrent single sends, so this is a no-op forworld-localandworld-postgres.Not in scope
This addresses the TTFS component. The larger TTLS tail at 256 branches is a separate problem and is not fixed here: in an isolated run, all 256 steps were served by only 10 compute instances (max 82 in flight), and per-step wall time inflated from ~200 ms to 1,100–2,400 ms under that packing. That is what drives TTLS to ~14 s, and it needs its own investigation.
Separately, much of the published durabench 128-branch number is a harness artifact — the sweep runs cells concurrently, so 3,000–4,000 branch steps from the 512/1024 cells were in flight during the 128/256 measurements. Isolated, 128-branch TTLS is ~1.5 s rather than the charted 9.3 s.
Testing
packages/world-vercel/src/queue.test.ts— 6 new cases: one request per fan-out with input order and per-message idempotency keys preserved; splitting at the 100 cap; per-entry failure reported without rejecting; a short result array flagged retryable; separate requests per region; empty input.packages/core/src/runtime/helpers.test.ts— 6 new cases: batch used when available; fallback to single sends; rejection naming the shortfall;retryablepropagation; deferred treated as success; empty input.core'squickjs-runtime.test.tsfails to importquickjs-assets.generated.jsboth with and without this change — pre-existing, verified against a clean tree.)Follow-ups
runtime.ts, backstop/recovery wakes) still publishes per step. It is skipped on the clean fan-out path, so it isn't hot, but it could usequeueMessagestoo.WORKFLOW_RESILIENT_STEP_DISPATCH, off by default) pairs each create with its own publish and is mutually exclusive with the batched fold, so it does not benefit.🤖 Generated with Claude Code