Skip to content

perf(world-vercel): batch a fan-out's step-execution queue publishes - #3838

Merged
karthikscale3 merged 4 commits into
mainfrom
pgp/exponential-ttls-fix
Sep 11, 2026
Merged

karthikscale3 merged 4 commits into
mainfrom
pgp/exponential-ttls-fix

Conversation

@pranaygp

Copy link
Copy Markdown
Contributor

Why

Investigating a durabench report that TTLS (time-to-last-step) grows sharply past 64 parallel branches, I measured where a fan-out's time actually goes. Dispatch is not the problem — all 256 step_created events land within one second, because that path is already batched (createBatch, 32-event chunks) and HTTP/2-multiplexed.

The queue publishes were not. The batched fan-out fold published one message per step:

await Promise.all(stepEntries.map((entry) => queueMessage(world, …)))

Those sends deliberately ride the shared default undici agent — connections: 8, pipelining: 1, allowH2: false (see getQueueDispatcher) — so an N-branch fan-out costs roughly N/8 serialized HTTP/1.1 round trips. And handleSuspension is awaited in full before any inline step body runs (runtime.ts:3212 → 4126), so that cost lands directly on time-to-first-step rather than overlapping with useful work.

Measured isolated TTFS medians on the current beta, one run at a time on an idle deployment: 64 branches 505 ms, 128 branches 633 ms, 256 branches 785 ms.

What

  • @vercel/queue 0.5.0 → 0.5.1, which adds experimental_sendBatch (up to 100 messages per request, per-entry results).
  • Queue.queueBatch (optional) on @workflow/world. Worlds that don't implement it are unaffected.
  • queueBatch in @workflow/world-vercel, on experimental_sendBatch. It groups messages by the routing dimensions one VQS request cannot span (region / deploymentId / transport / physical topic) and splits at the 100-message cap; neither split is observable in the returned order. For the case this exists for — one run's fan-out to one logical queue — that is a single group.
  • queueMessages in @workflow/core, used by the fan-out fold's publishChunkSteps. Each commit chunk now publishes in one request instead of up to 32.

Error semantics

queueBatch reports per-entry outcomes rather than throwing, because a batch can partially fail; it rejects only for a request-level failure where no entry outcome is known.

queueMessages keeps this call site's existing all-or-nothing behavior: it rejects if any entry failed, so the delivery is redelivered and republishes the whole set. That is safe precisely because the caller already passes a per-step idempotencyKey (stepDispatchIdempotencyKey) — without it, republishing would redeliver the steps that already succeeded. The interface docs make that a requirement rather than a suggestion.

messageId: null with no error is an acceptance (deferred, no ID yet), matching queue(), so success is error === undefined rather than a non-null id.

Worlds without queueBatch fall back to concurrent single sends, so this is a no-op for world-local and world-postgres.

Not in scope

This addresses the TTFS component. The larger TTLS tail at 256 branches is a separate problem and is not fixed here: in an isolated run, all 256 steps were served by only 10 compute instances (max 82 in flight), and per-step wall time inflated from ~200 ms to 1,100–2,400 ms under that packing. That is what drives TTLS to ~14 s, and it needs its own investigation.

Separately, much of the published durabench 128-branch number is a harness artifact — the sweep runs cells concurrently, so 3,000–4,000 branch steps from the 512/1024 cells were in flight during the 128/256 measurements. Isolated, 128-branch TTLS is ~1.5 s rather than the charted 9.3 s.

Testing

  • packages/world-vercel/src/queue.test.ts — 6 new cases: one request per fan-out with input order and per-message idempotency keys preserved; splitting at the 100 cap; per-entry failure reported without rejecting; a short result array flagged retryable; separate requests per region; empty input.
  • packages/core/src/runtime/helpers.test.ts — 6 new cases: batch used when available; fallback to single sends; rejection naming the shortfall; retryable propagation; deferred treated as success; empty input.
  • Full suites green: core 2,252 passed, world-vercel 567 passed, world 165 passed. (core's quickjs-runtime.test.ts fails to import quickjs-assets.generated.js both with and without this change — pre-existing, verified against a clean tree.)

Follow-ups

  • The runtime's own dispatch pass (runtime.ts, backstop/recovery wakes) still publishes per step. It is skipped on the clean fan-out path, so it isn't hot, but it could use queueMessages too.
  • The resilient-dispatch path (WORKFLOW_RESILIENT_STEP_DISPATCH, off by default) pairs each create with its own publish and is mutually exclusive with the batched fold, so it does not benefit.

🤖 Generated with Claude Code

A `Promise.all` fan-out dispatched one queue message per branch. Those
publishes ride the shared default undici agent (8 connections, HTTP/1.1,
`pipelining: 1` — see `getQueueDispatcher`), and `handleSuspension` is
awaited in full before the first inline step body runs, so an N-branch
fan-out paid ~N/8 serialized round trips straight onto time-to-first-step.
The `step_created` writes were already batched and HTTP/2-multiplexed; the
publishes were the remaining per-branch round trip.

Adds an optional `Queue.queueBatch`, implemented on `@vercel/queue`'s
`experimental_sendBatch` (0.5.1), and uses it for the batched fan-out fold's
publishes. Each commit chunk now publishes in one request instead of up to
32.

`queueBatch` reports per-entry outcomes rather than throwing, because a
batch can partially fail. `queueMessages` in core keeps the previous
all-or-nothing behavior for this call site: it rejects if any entry failed,
so the delivery is redelivered and republishes the set, deduped by the
per-step `idempotencyKey` the caller already passed. Worlds without
`queueBatch` fall back to concurrent single sends.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
example-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workflow-docs Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC
workflow-web Ready Ready Preview, v0 Sep 11, 2026 5:50pm UTC

@changeset-bot

changeset-bot Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 367e5e0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
Name Type
@workflow/world-vercel Patch
@workflow/world Patch
@workflow/core Patch
@workflow/cli Patch
@workflow/web Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/world-local Patch
@workflow/world-postgres Patch
@workflow/world-testing Patch
@workflow/builders Patch
@workflow/next Patch
@workflow/nitro Patch
workflow Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • addTenWorkflow (express)
  • addTenWorkflow (nitro)
  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (sveltekit)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • cold-start-warmup · suite warmup (astro) · at 17:50:11Z · abandoned wrun_01M28SE5C4ZEPN603PKE3RH2VN
  • cold-start-warmup · suite warmup (tanstack-start) · at 17:50:17Z · abandoned wrun_01M28SE77S9FWMAG6XH0XTN056
  • run-pickup-stall · cross-file step error preserves message and function names in stack (nextjs-webpack) · at 17:57:15Z · abandoned wrun_01M28SV74WYVD67JR27872GV7F
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 17:57:23Z · abandoned wrun_01M28SVE8HV5VZM1AQTHAMAQBW
  • run-pickup-stall · concurrent hook token conflict - two workflows cannot use the same hook token simultaneously (nextjs-webpack) · at 17:57:26Z · abandoned wrun_01M28SVHJ86R19WV3N2M60VXJ3

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3662 0 685 4347
✅ 💻 Local Development 3922 0 586 4508
✅ 📦 Local Production 3922 0 586 4508
✅ 🐘 Local Postgres 3922 0 586 4508
✅ 🪟 Windows 320 0 2 322
✅ 🌐 Cross-language Conformance 68 0 74 142
✅ vercel-http-transport 823 0 143 966
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 557 0 87 644
Total 17223 0 2749 19972
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 133 0 28
✅ astro-quickjs 133 0 28
✅ example-node 133 0 28
✅ example-quickjs 133 0 28
✅ express-node 133 0 28
✅ express-quickjs 133 0 28
✅ fastify-node 133 0 28
✅ fastify-quickjs 133 0 28
✅ hono-node 133 0 28
✅ hono-quickjs 133 0 28
✅ nest-node 133 0 28
✅ nest-quickjs 133 0 28
✅ nextjs-turbopack-node 158 0 3
✅ nextjs-turbopack-quickjs 158 0 3
✅ nextjs-webpack-node 158 0 3
✅ nextjs-webpack-quickjs 158 0 3
✅ nitro-node 133 0 28
✅ nitro-quickjs 133 0 28
✅ nuxt-node 133 0 28
✅ nuxt-quickjs 133 0 28
✅ python-node 66 0 95
✅ sveltekit-node 152 0 9
✅ sveltekit-quickjs 152 0 9
✅ tanstack-start-node 133 0 28
✅ tanstack-start-quickjs 133 0 28
✅ vite-node 133 0 28
✅ vite-quickjs 133 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 141 0 20
✅ nextjs-turbopack-canary-quickjs 141 0 20
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 141 0 20
✅ nextjs-webpack-canary-quickjs 141 0 20
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 141 0 20
✅ nextjs-turbopack-canary-quickjs 141 0 20
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 141 0 20
✅ nextjs-webpack-canary-quickjs 141 0 20
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 141 0 20
✅ nextjs-turbopack-canary-quickjs 141 0 20
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 141 0 20
✅ nextjs-webpack-canary-quickjs 141 0 20
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 160 0 1
✅ nextjs-turbopack-quickjs 160 0 1

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 74

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 133 0 28
✅ express 133 0 28
✅ hono 133 0 28
✅ nextjs-turbopack 158 0 3
✅ nitro 133 0 28
✅ vite 133 0 28

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 133 0 28
✅ express 133 0 28
✅ nextjs-turbopack 158 0 3
✅ vite 133 0 28

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 367e5e0 · Fri, 11 Sep 2026 18:11:16 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1600 (+448%) 🔻 1742 🔴 (+31%) 🔻 1781 🔴 (+25%) 🔻 1937 🔴 (+10.0%) 30
TTFS stream 164 (-87%) 💚 1721 🔴 (+33%) 🔻 1760 🔴 (+34%) 🔻 2126 🔴 (+58%) 🔻 30
TTFS hook + stream 642 (-59%) 💚 2164 🔴 (+28%) 🔻 2247 🔴 (+25%) 🔻 6701 🔴 (+216%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 680 (+15%) 2612 (+240%) 🔻 2674 (+189%) 🔻 2860 (+44%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 2363 (+30%) 🔻 4485 (-1.3%) 5088 (+4.9%) 12571 (+69%) 🔻 10
STSO 1020 steps (inline) 110 (-17%) 💚 177 (+9.3%) 193 (+7.2%) 258 (+9.3%) 1019
WO 1020 steps 179119 (+9.1%) 179119 (+9.1%) 179119 (+9.1%) 179119 (+9.1%) 1
CRTT first chunk (pooled) 77 (+40%) 🔻 120 (+46%) 🔻 133 (+40%) 🔻 226 (+97%) 🔻 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 93.5 (+26%) 202 (+66%) 261 (+41%) 492 (+79%) 154 (+35%) 10
size sweep (100/s, 160B-12KB) 105 (+60%) 204 (+24%) 266 (-3%) 376 (-22%) 183 (-22%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 106 (+29%) 165 (+20%) 201 (-9%) 299 (-35%) 237 (+40%) 3
replay eve-gpt-5.6-sol-2000t (1x) 127 (+54%) 177 (+30%) 209 (-51%) 373 (-87%) 253 (-85%) 2
replay eve-gpt-5.6-sol-2000t (2x) 89 (+13%) 290 (+14%) 436 (-81%) 682 (-81%) 257 (-42%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 163045ms → this run 177564ms (Δ +14519ms, +9%)

100-150 ms  ┃█████                    main 247  this   1  -246
150-200 ms  ███████████████████░░░░┃  main 741  this 947  +206
200-250 ms  █┃                        main  21  this  60   +39
250-300 ms  ┃                         main   5  this   6    +1
300-350 ms  ┃                         main   1  this   3    +2
350-400 ms  ┃                         main   4  this   0    -4
400-450 ms  ┃                         main   0  this   1    +1
450-500 ms  ┃                         main   0  this   1    +1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▁█▂▁···  159.6 (+44%)  151 (+45%)  261 (+41%)  492 (+79%)  3000
sweep    ······▁█▃····  172.3 (+31%)  160 (+45%)   266 (-3%)  376 (-22%)  3000
gw 1x    ······▂█▁····  143.1 (+25%)  138 (+31%)   201 (-9%)  299 (-35%)  5295
eve 1x   ·····▁▂█▂▁···  146.8 (-24%)  133 (+37%)  209 (-51%)  373 (-87%)  5186
eve 2x   ·····▁▁█▆▁···  213.3 (-26%)  190 (+45%)  436 (-81%)  682 (-81%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  █▅▇▅▅▃▃▂▁▃  133–191ms
sweep    █▇▅█▅▅▂▃▄▁  147–195ms
gw 1x    █▄▆▄▃▃▃▁▆▅  127–159ms
eve 1x   ▄▄▅▆▅▄██▄▁  124–166ms
eve 2x   ▆▂▁▁▂▃▆█▆▃  166–289ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▂██▆▅▁▃  170–175ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  █▅▆▆▃▄▃▃▁▆  37–61ms
sweep    ▆▇██▄▆▂▅▃▁  57–76ms
gw 1x    █▆▂▁▂▄▄▂▅█  40–52ms
eve 1x   ▇▄▅█▅▅▇▅▁▆  20–31ms
eve 2x   █▆▁▂▄▃▄▅▃▅  21–39ms
📜 Previous results (2)

480b34f

Fri, 11 Sep 2026 17:11:54 GMT · run logs

vercel / nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1537 (+554%) 🔻 1772 🔴 (+35%) 🔻 1806 🔴 (+34%) 🔻 2014 🔴 (+17%) 🔻 30
TTFS stream 160 (-27%) 💚 1761 🔴 (+36%) 🔻 1782 🔴 (+36%) 🔻 1979 🔴 (+26%) 🔻 30
TTFS hook + stream 1915 (+27%) 🔻 2178 🔴 (+31%) 🔻 2212 🔴 (+31%) 🔻 2441 🔴 (+34%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 600 (-2.6%) 748 (-4.8%) 2066 (+4.0%) 2377 (+18%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 2392 (+28%) 🔻 5764 (+102%) 🔻 6521 (-12%) 9390 (+2.2%) 10
STSO 1020 steps (inline) 138 (+13%) 162 (+8.0%) 180 (+4.7%) 262 (+5.2%) 1019
WO 1020 steps 163501 (+7.3%) 163501 (+7.3%) 163501 (+7.3%) 163501 (+7.3%) 1
CRTT first chunk (pooled) 56 (+9.8%) 84 (+3.7%) 105 (-33%) 💚 209 (+27%) 🔻 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 77.5 (+10%) 146 (-8%) 254 (-18%) 516 (-11%) 122 (-8%) 10
size sweep (100/s, 160B-12KB) 74 (-1%) 127 (-16%) 260 (+20%) 538 (-26%) 101 (-45%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 98 (+36%) 114 (-13%) 141 (-18%) 382 (-32%) 249 (-16%) 3
replay eve-gpt-5.6-sol-2000t (1x) 78.5 (+5%) 116 (-9%) 156 (-14%) 410 (+1%) 378 (+33%) 2
replay eve-gpt-5.6-sol-2000t (2x) 84 (-9%) 170 (-12%) 221 (-21%) 562 (+13%) 248 (-12%) 3

68b61ac

Thu, 27 Aug 2026 00:26:58 GMT · run logs

vercel / nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 174 (-64%) 💚 1268 🔴 (+18%) 🔻 1291 🔴 (+15%) 🔻 1412 🔴 (-4.3%) 30
TTFS stream 188 (-80%) 💚 1239 🔴 (+18%) 🔻 1287 🔴 (+19%) 🔻 1374 🔴 (+2.8%) 30
TTFS hook + stream 1282 (+187%) 🔻 1583 🔴 (+22%) 🔻 1708 🔴 (+20%) 🔻 1810 🔴 (-63%) 💚 30
Fan-out TTFS Promise.all(100 steps) 520 (+8.3%) 768 (+6.1%) 1703 (+118%) 🔻 1829 (-63%) 💚 10
Fan-out TTLS Promise.all(100 steps) 2539 (+47%) 🔻 8826 (+198%) 🔻 9262 (+93%) 🔻 9473 (-9.5%) 10
STSO 1020 steps (inline) 71 (-34%) 💚 153 (+16%) 🔻 177 (+18%) 🔻 286 (+42%) 🔻 1019
WO 1020 steps 150717 (+12%) 150717 (+12%) 150717 (+12%) 150717 (+12%) 1
CRTT first chunk (pooled) 93 (+5.7%) 136 (+4.6%) 206 (+1.5%) 417 (+77%) 🔻 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 133 (+18%) 157 (+7%) 207 (-27%) 370 (-52%) 143 (-34%) 10
size sweep (100/s, 160B-12KB) 115 (+1%) 175 (+7%) 218 (-35%) 373 (-36%) 151 (-38%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 119 (-1%) 144 (+14%) 181 (+3%) 277 (-50%) 224 (-40%) 3
replay eve-gpt-5.6-sol-2000t (1x) 255 (+53%) 181 (+24%) 547 (+159%) 4507 (+606%) 2291 (+336%) 2
replay eve-gpt-5.6-sol-2000t (2x) 122 (-10%) 216 (+13%) 315 (+8%) 541 (-23%) 320 (+12%) 3
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@socket-security

socket-security Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Updatednpm/​@​vercel/​queue@​0.5.0 ⏵ 0.5.196100100 +198 +2100

View full report

@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 236.0 KiB (±0) 42.3 KiB (±0) 1.81 MiB (+2.4 KiB)
nextjs-turbopack 242.2 KiB (±0) 439 B (±0) 837.4 KiB (+973 B)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

367e5e0 · run

@karthikscale3 karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Reviewed the change and load-tested the batched path at 64 branches. Design is sound and the error semantics are the right ones — per-entry outcomes rather than throws, rejection reserved for request-level failures, idempotency keys documented as a requirement rather than a suggestion. One thing I'd fix before merge, the rest are notes.

Verified independently: @vercel/queue 0.5.0 → 0.5.1 is purely additive (no changes to the single-send path); createQueue's result is spread into the World so queueBatch survives composition; the grouping key covers every routing dimension prepareSend derives; and the prepareSend/clientFor refactor is behavior-preserving for queue().

Wire-level, against a fake VQS over real undici (real createQueue, real SDK, real multipart — nothing mocked below createQueue): 64 branches go out in one request, keys and order intact on the wire; per-entry 429/400 are reported without rejecting; request-level 500 rejects; 140 messages split [100, 40] with global order preserved; and a full republish after a partial failure dedupes 61 and delivers 64 unique. The proxy path handles the /batch suffix and the chunked multipart body fine (client side — this says nothing about api.vercel.com's own allowlist, worth confirming separately if a proxy-mode caller ever adopts queueBatch).

Wedge hunt at the runtime level, driving the real fold with a redelivery loop and a VQS dedupe model, every scenario also run against a world without queueBatch as a control:

batched control (pre-PR path)
happy 64 63 published in 2 requests [32,31], 1 delivery 63 requests
transient entry failures converges in 2, 0 duplicate steps same
permanent entry failure never converges, 96 requests never converges, 3024 requests
flaky request-level 503 converges in 2 same
3x total outage converges in 4, events committed once same

No wedge is introduced by batching — every failure mode I could reach has an identical outcome on the control, and in the bad case batching is strictly better (96 requests vs 3024 to reach the same dead end). Zero duplicate step deliveries anywhere; the idempotency keys do their job. deferredBatchWork correctly surfaces a failed batched publish rather than swallowing a chunk.

One heads-up for whoever owns the dashboards: fan-out queue.publish span volume drops ~32x and rpc.method becomes publishBatch. messaging.batch.message_count recovers the count, which is the right call, but any monitor counting those spans will see a step change.

},
async () => {
const results = await batch(queueName, messages);
const failures = results.filter((result) => result.error !== undefined);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: queueMessages only inspects error, never results.length. A World whose queueBatch returns a short array reports success here: handleSuspension resolves, the delivery is acked, and those steps are never dispatched — the run wedges with no error anywhere in the system.

world-vercel guards this internally (toBatchResult(undefined)) and the SDK length-checks too, so it's unreachable through the world in this PR. But queueBatch is documented in building-a-world.mdx for third-party worlds, which makes it an unenforced contract at exactly the boundary that publishes it — and unlike a rejection, this failure is silent.

I reproduced it with a 64-branch fan-out against a world returning half the results: 32 of 63 steps silently lost, handleSuspension resolved without error. Adding the length check flips it to a clean rejection that converges on redelivery, and all 143 tests in helpers.test.ts + suspension-handler.test.ts still pass:

const results = await batch(queueName, messages);
if (results.length !== messages.length) {
  throw Object.assign(
    new Error(
      `Queue batch for ${queueName} returned ${results.length} result(s) ` +
        `for ${messages.length} message(s)`
    ),
    { retryable: true }
  );
}

);
// Surfaced so a caller (and the delivery-level retry above it) can tell
// a transient partial batch from a permanent rejection.
Object.assign(error, { retryable });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: retryable is never read anywhere in packages/core/src/runtime, so the comment above it overstates what happens: queueMessages throws identically for retryable and non-retryable, and the delivery retries either way.

That has a measurable cost. With one permanently-rejected entry in a 64-branch fan-out, the delivery burns all 48 redeliveries and ~2,900 redundant republishes before giving up. The unbatched path does the same, so this isn't a regression — but the batched path is the one that now has the information and discards it, which makes a fast-fail on retryable: false cheap to add later.

>();
for (const [index, entry] of messages.entries()) {
const prepared = prepareSend(queueName, entry.message, entry.opts);
const key = `${prepared.region}${prepared.deploymentId}${prepared.useCbor}${prepared.topic}`;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: This grouping is a no-op under WORKFLOW_SEQUENTIAL_REPLAYS=1. stepDispatch.queueName is the flow topic, which matches FLOW_TOPIC_PATTERN, so getPhysicalQueueName returns ${queue}_${runId}_${stepId} — a distinct topic per message, hence a distinct group per message, hence N requests of one message each.

Off by default so there's no production impact, but two consequences worth a line of comment: the win silently disappears for opted-in users, and each send now goes through the multipart batch endpoint rather than send(), which loses the SDK's 502 consumer_discovery_failed → ConsumerDiscoveryError mapping (sendBatch only maps 503). A client.send fast path for single-entry groups would restore both.

Comment thread packages/world-vercel/src/queue.test.ts Outdated
process.env.VERCEL_DEPLOYMENT_ID = 'dpl_batch';
});
afterEach(() => {
process.env.VERCEL_DEPLOYMENT_ID = undefined;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Nit: process.env.VERCEL_DEPLOYMENT_ID = undefined sets the string "undefined"; the rest of the file uses delete. Harmless as the last describe in the file, just inconsistent.

@pranaygp
pranaygp marked this pull request as ready for review September 11, 2026 07:17
@pranaygp
pranaygp requested a review from a team as a code owner September 11, 2026 07:17
Copilot AI lite review requested due to automatic review settings September 11, 2026 07:17
`experimental_sendBatch` injects the active trace context into the multipart
REQUEST headers, and the per-part headers it builds never see it. VQS stores
headers per message and re-emits a stored `traceparent` at delivery as
`x-vercel-queue-traceparent`, which is what lets a consumer attach a span link
back to its producer, so a batched message arrived with no producer context
and its `vqs.process` span got no link. `send()` is unaffected: for a single
message the request headers ARE that message's headers.

At 64 branches that was 63 of 64 step dispatches losing the transport-level
producer link. The run's own step tracing was never affected: that carrier
travels in the message payload (`WorkflowInvokePayload.traceCarrier`), which
is what the consumer builds its trace context from, not a header.

Injects the active context into each entry's headers in `queueBatch` — last,
so it wins over caller-supplied `opts.headers` exactly as the SDK's own
injection does — and honors VERCEL_QUEUE_TRACE_PROPAGATION so that kill
switch still covers both paths. `getTraceContextHeaders()` is factored out of
`injectTraceContextIntoHeaders` so the two share one source.

Verified on the wire against a stub VQS speaking the real batch endpoint:
`traceparent` carrying the producer's traceId/spanId lands on all 64
multipart parts through the real SDK, with the per-message idempotency keys
still alongside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@karthikscale3
karthikscale3 merged commit e00b1a5 into main Sep 11, 2026
185 checks passed
@karthikscale3
karthikscale3 deleted the pgp/exponential-ttls-fix branch September 11, 2026 18:23
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for e00b1a5 (AI decision).

This is a performance optimization that adds new API surface: an optional Queue.queueBatch method on @workflow/world, its @vercel/queue experimental_sendBatch implementation in @workflow/world-vercel, a new queueMessages helper in core, and a dependency bump to obtain the batch API. Fan-out dispatch on stable is functionally correct today, just slower, so this does not qualify as a stability fix. The two fix(...) commits in the series (short-result rejection and per-message trace context) only harden code introduced by this same PR and have nothing to correct on stable.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

e00b1a57ee8e4cc7b597e1bd23188efe23cc6235

This branch was successfully deployed

18 active deployments
Preview – workflow-swc-playground — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workflow-docs — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-astro-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-vite-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – example-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workflow-tarballs — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-express-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-hono-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workflow-web — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Preview – workbench-python-workflow — 367e5e09 Deployed Sep 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants