Skip to content

fix(world-vercel): split large frames on the WebSocket events transport - #4504

Merged
shalabhc merged 9 commits into
mainfrom
switch-ai-gateway-transport-from-http-to-websock-split-frames
Sep 30, 2026
Merged

shalabhc merged 9 commits into
mainfrom
switch-ai-gateway-transport-from-http-to-websock-split-frames

Conversation

@shalabhc

@shalabhc shalabhc commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Why

Some WebSocket paths limit the size of a single message, commonly to 16 MiB (2^24 bytes). Event payloads can be bigger than that: a step's input travels inside the step_started reply, and its output inside the step_completed request. On such a path an oversized message never arrives and the socket is torn down, while the same write over HTTP works. This change makes the WebSocket events transport usable on those paths.

What

Any frame whose encoded size is over a message limit is sent as several WebSocket messages, and the receiver rebuilds it before handling it. Frames at or under the limit go out exactly as before.

Meta Body
Whole frame (≤ limit) the frame's meta, unchanged whole body
First part the frame's meta + partIndex: 0, partCount first piece of the body
Continuation part exactly { type: 'part', reqId, partIndex, partCount } next piece of the body

Compatibility:

  • Opt-in for replies. Every upgrade request carries x-workflow-ws-flags: frame-parts. The backend splits replies only for a client that lists the flag; clients without it keep receiving whole replies. This is the same shape as the Accept opt-in the hook_received preload uses: the client names what it understands, and the backend keeps the old behaviour otherwise. The header is a list, so later protocol flags can be added the same way.
  • Requests. The client splits any request over the limit. This needs a backend that accepts split frames, so the backend has to support them before this is released.

Behaviour:

  • Limit: WORKFLOW_WS_MAX_MESSAGE_BYTES, default 12 MiB. Read with envNumber, clamped to 2 MiB–16 MiB with a one-time warning. It bounds every message, header included.
  • Which frames split: only frames with a reqId, which is how continuations are matched to their frame.
  • Reassembly: keyed by reqId. Whole frames for other requests may arrive between two parts.
  • Settled requests: a split reply for a request that has already settled is read through, not buffered, and then dropped with a log line, like any late reply.
  • Order: replies are decoded synchronously, so parts are assembled in arrival order.
  • Protocol errors: a part that breaks the protocol fails the connection, the same as an undecodable frame. That covers an orphan continuation, a skipped or repeated index, a changed count, extra fields, and a whole reply reusing the reqId of an open split reply.
  • Limits (the same as the backend's): at most one open split frame, of at most 256 MiB, and at most 257 parts per frame, of any size. At the 2 MiB floor of the message limit, 257 parts still carry 256 MiB. A request whose body is over 256 MiB is refused before anything is sent (WsFrameTooLargeError).

Changes:

  • ws-parts.ts (new):
    • the flag header and constants;
    • encodeWsFrameMessages / splitEncodedFrame;
    • WsPartAssembler.
  • frames.ts: a synchronous single-frame decodeFrame.
  • ws-transport.ts:
    • offers the flag on every upgrade;
    • requests are split before sending;
    • every received message goes through a per-connection assembler.
  • Telemetry: workflow.events.ws.request_parts and workflow.events.ws.reply_parts on the write span, set only when a message was split.
  • Docs: WORKFLOW_WS_MAX_MESSAGE_BYTES in the configuration reference and the Vercel World page.

The transport stays opt-in (default http).

Tests

  • ws-parts.test.ts:
    • split/rebuild at and around the limit;
    • frames under the limit unchanged;
    • interleaving;
    • unwanted frames read through;
    • every protocol-error case, including reqId reuse;
    • each limit, including the encoder refusing more than 257 parts or a body over 256 MiB;
    • limit parsing and clamping;
    • a byte-for-byte golden fixture shared with the backend implementation.
  • ws-transport.test.ts:
    • the flag is offered on the upgrade;
    • a large request is sent as parts under the limit that rebuild to the frame;
    • a small request stays one message;
    • a split reply resolves the right request with another reply between its parts;
    • a send error on a later part fails the request;
    • a close partway through a split reply fails the request;
    • a split reply for a settled request is read through;
    • a bad part fails the connection.
  • tsc --noEmit and the full @workflow/world-vercel suite pass (824 tests).

@shalabhc
shalabhc requested a review from a team as a code owner September 30, 2026 03:28
@changeset-bot

changeset-bot Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 70520fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
Name Type
@workflow/world-vercel Patch
@workflow/cli Patch
@workflow/core Patch
@workflow/web Patch
workflow Patch
@workflow/world-testing Patch
@workflow/builders Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
example-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workflow-docs Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC
workflow-web Ready Ready Preview, v0 Sep 30, 2026 9:34pm UTC

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • 'hookGetConflictWithPriorStepWorkflow' - hook.getConflict() does not block step execution (nextjs-webpack · local-dev / local / node / stable)
  • hookGetConflictWorkflow - awaiting hook.getConflict() registers hook without payload (nextjs-webpack · local-dev / local / node / stable)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · runs app-generated source against a registered step (tanstack-start) · at 21:36:56Z · abandoned wrun_01M3T3Z49SWVH9JJA42DX1A73R
  • cold-start-warmup · suite warmup (tanstack-start) · at 21:37:04Z · abandoned wrun_01M3T3Z49S4HG1KSQK9DDKJ2KQ
  • run-pickup-stall · step-return-value serialization failure is catchable in workflow code (nuxt) · at 21:37:39Z · abandoned wrun_01M3T40DW16XDZ7MW338J9ZFXG
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 21:41:38Z · abandoned wrun_01M3T47Q9ACGZKHKAMFQFCN42Y
  • run-pickup-stall · hookGetConflictWorkflow - awaiting hook.getConflict() registers hook without payload (nextjs-webpack) · at 21:41:48Z · abandoned wrun_01M3T481DSNK36R5S9Q4660H98
  • run-pickup-stall · 'hookGetConflictWithPriorStepWorkflow' - hook.getConflict() does not block step execution (nextjs-webpack) · at 21:41:48Z · abandoned wrun_01M3T481J7NVYXM28P6YFQDGEG
  • run-pickup-stall · ChainableService.processWithThis - static step methods using this to reference the class (nextjs-webpack) · at 21:43:03Z · abandoned wrun_01M3T4AA7WG19WRGM4J2QX7YG3
  • run-pickup-stall · thisSerializationWorkflow - step function invoked with .call() and .apply() (nextjs-webpack) · at 21:43:04Z · abandoned wrun_01M3T4ABM0882JM69ZQSN0YWZ9

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3904 0 875 4779
✅ 💻 Local Development 4406 0 550 4956
✅ 📦 Local Production 4406 0 550 4956
✅ 🐘 Local Postgres 4406 0 550 4956
✅ 🪟 Windows 342 0 12 354
✅ 🌐 Cross-language Conformance 68 0 84 152
✅ dynamic-runs 0 0 0 0
✅ vercel-http-transport 879 0 183 1062
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 595 0 113 708
Total 19033 0 2917 21950
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 142 0 35
✅ astro-quickjs 142 0 35
✅ example-node 142 0 35
✅ example-quickjs 142 0 35
✅ express-node 142 0 35
✅ express-quickjs 142 0 35
✅ fastify-node 142 0 35
✅ fastify-quickjs 142 0 35
✅ hono-node 142 0 35
✅ hono-quickjs 142 0 35
✅ nest-node 142 0 35
✅ nest-quickjs 142 0 35
✅ nextjs-turbopack-node 169 0 8
✅ nextjs-turbopack-quickjs 169 0 8
✅ nextjs-webpack-node 169 0 8
✅ nextjs-webpack-quickjs 169 0 8
✅ nitro-node 142 0 35
✅ nitro-quickjs 142 0 35
✅ nuxt-node 142 0 35
✅ nuxt-quickjs 142 0 35
✅ python-node 66 0 111
✅ sveltekit-node 161 0 16
✅ sveltekit-quickjs 161 0 16
✅ tanstack-start-node 142 0 35
✅ tanstack-start-quickjs 142 0 35
✅ vite-node 142 0 35
✅ vite-quickjs 142 0 35

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 148 0 29
✅ astro-stable-quickjs 148 0 29
✅ express-stable-node 148 0 29
✅ express-stable-quickjs 148 0 29
✅ fastify-stable-node 148 0 29
✅ fastify-stable-quickjs 148 0 29
✅ hono-stable-node 148 0 29
✅ hono-stable-quickjs 148 0 29
✅ nest-stable-node 148 0 29
✅ nest-stable-quickjs 148 0 29
✅ nextjs-turbopack-canary-node 176 0 1
✅ nextjs-turbopack-canary-quickjs 176 0 1
✅ nextjs-turbopack-stable-node 176 0 1
✅ nextjs-turbopack-stable-quickjs 176 0 1
✅ nextjs-webpack-canary-node 176 0 1
✅ nextjs-webpack-canary-quickjs 176 0 1
✅ nextjs-webpack-stable-node 176 0 1
✅ nextjs-webpack-stable-quickjs 176 0 1
✅ nitro-stable-node 148 0 29
✅ nitro-stable-quickjs 148 0 29
✅ nuxt-stable-node 148 0 29
✅ nuxt-stable-quickjs 148 0 29
✅ sveltekit-stable-node 167 0 10
✅ sveltekit-stable-quickjs 167 0 10
✅ tanstack-start-node 148 0 29
✅ tanstack-start-quickjs 148 0 29
✅ vite-stable-node 148 0 29
✅ vite-stable-quickjs 148 0 29

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 148 0 29
✅ astro-stable-quickjs 148 0 29
✅ express-stable-node 148 0 29
✅ express-stable-quickjs 148 0 29
✅ fastify-stable-node 148 0 29
✅ fastify-stable-quickjs 148 0 29
✅ hono-stable-node 148 0 29
✅ hono-stable-quickjs 148 0 29
✅ nest-stable-node 148 0 29
✅ nest-stable-quickjs 148 0 29
✅ nextjs-turbopack-canary-node 176 0 1
✅ nextjs-turbopack-canary-quickjs 176 0 1
✅ nextjs-turbopack-stable-node 176 0 1
✅ nextjs-turbopack-stable-quickjs 176 0 1
✅ nextjs-webpack-canary-node 176 0 1
✅ nextjs-webpack-canary-quickjs 176 0 1
✅ nextjs-webpack-stable-node 176 0 1
✅ nextjs-webpack-stable-quickjs 176 0 1
✅ nitro-stable-node 148 0 29
✅ nitro-stable-quickjs 148 0 29
✅ nuxt-stable-node 148 0 29
✅ nuxt-stable-quickjs 148 0 29
✅ sveltekit-stable-node 167 0 10
✅ sveltekit-stable-quickjs 167 0 10
✅ tanstack-start-node 148 0 29
✅ tanstack-start-quickjs 148 0 29
✅ vite-stable-node 148 0 29
✅ vite-stable-quickjs 148 0 29

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 148 0 29
✅ astro-stable-quickjs 148 0 29
✅ express-stable-node 148 0 29
✅ express-stable-quickjs 148 0 29
✅ fastify-stable-node 148 0 29
✅ fastify-stable-quickjs 148 0 29
✅ hono-stable-node 148 0 29
✅ hono-stable-quickjs 148 0 29
✅ nest-stable-node 148 0 29
✅ nest-stable-quickjs 148 0 29
✅ nextjs-turbopack-canary-node 176 0 1
✅ nextjs-turbopack-canary-quickjs 176 0 1
✅ nextjs-turbopack-stable-node 176 0 1
✅ nextjs-turbopack-stable-quickjs 176 0 1
✅ nextjs-webpack-canary-node 176 0 1
✅ nextjs-webpack-canary-quickjs 176 0 1
✅ nextjs-webpack-stable-node 176 0 1
✅ nextjs-webpack-stable-quickjs 176 0 1
✅ nitro-stable-node 148 0 29
✅ nitro-stable-quickjs 148 0 29
✅ nuxt-stable-node 148 0 29
✅ nuxt-stable-quickjs 148 0 29
✅ sveltekit-stable-node 167 0 10
✅ sveltekit-stable-quickjs 167 0 10
✅ tanstack-start-node 148 0 29
✅ tanstack-start-quickjs 148 0 29
✅ vite-stable-node 148 0 29
✅ vite-stable-quickjs 148 0 29

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 171 0 6
✅ nextjs-turbopack-quickjs 171 0 6

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 84

✅ dynamic-runs

App Passed Failed Skipped
✅ astro-vercel 0 0 0
✅ example-vercel 0 0 0
✅ express-vercel 0 0 0
✅ fastify-vercel 0 0 0
✅ hono-vercel 0 0 0
✅ nest-vercel 0 0 0
✅ nextjs-turbopack-vercel 0 0 0
✅ nextjs-webpack-vercel 0 0 0
✅ nitro-vercel 0 0 0
✅ nuxt-vercel 0 0 0
✅ sveltekit-vercel 0 0 0
✅ tanstack-start-vercel 0 0 0
✅ vite-vercel 0 0 0

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 142 0 35
✅ express 142 0 35
✅ hono 142 0 35
✅ nextjs-turbopack 169 0 8
✅ nitro 142 0 35
✅ vite 142 0 35

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 142 0 35
✅ express 142 0 35
✅ nextjs-turbopack 169 0 8
✅ vite 142 0 35

📋 View full workflow run

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 70520fb · Wed, 30 Sep 2026 21:58:18 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 376 (-72%) 💚 2946 🔴 (+101%) 🔻 3070 🔴 (+101%) 🔻 3270 🔴 (+110%) 🔻 30
TTFS stream 379 (+137%) 🔻 2541 🔴 (+78%) 🔻 2703 🔴 (+80%) 🔻 2990 🔴 (+88%) 🔻 30
TTFS hook + stream 612 (-13%) 2813 🔴 (+46%) 🔻 2845 🔴 (+47%) 🔻 2985 🔴 (+42%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 591 (+22%) 🔻 3382 (+376%) 🔻 3487 (+282%) 🔻 3555 (+82%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 1541 (-20%) 💚 7328 (+198%) 🔻 8041 (+210%) 🔻 8761 (±0%) 10
STSO 1020 steps (inline) 124 (-17%) 💚 200 (+4.7%) 232 (+7.9%) 343 (+12%) 1019
WO 1020 steps 195154 (+2.5%) 195154 (+2.5%) 195154 (+2.5%) 195154 (+2.5%) 1
CRTT first chunk (pooled) 72 (-17%) 💚 110 (-19%) 💚 150 (-15%) 💚 271 (+33%) 🔻 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 97.5 (-8%) 159 (-39%) 242 (-26%) 396 (-9%) 146 (-32%) 10
size sweep (100/s, 160B-12KB) 94.5 (-26%) 157 (-66%) 454 (-35%) 680 (-54%) 128 (-50%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 89 (-23%) 149 (-43%) 175 (-55%) 323 (-49%) 198 (-53%) 3
replay eve-gpt-5.6-sol-2000t (1x) 182 (+32%) 176 (-36%) 271 (-57%) 3227 (+7%) 2246 (+32%) 2
replay eve-gpt-5.6-sol-2000t (2x) 109 (-28%) 204 (-66%) 320 (-66%) 505 (-81%) 195 (-73%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 190196ms → this run 194743ms (Δ +4547ms, +2%)

100-150 ms  ┃                         main   1  this  19   +18
150-200 ms  ████████████████████┃███  main 828  this 722  -106
200-250 ms  █████┃                    main 160  this 220   +60
250-300 ms  ┃                         main  18  this  38   +20
300-350 ms  ┃                         main   5  this  10    +5
350-400 ms  ┃                         main   3  this   3    +0
400-450 ms  ┃                         main   1  this   4    +3
450-500 ms  ┃                         main   1  this   2    +1
500-550 ms  ┃                         main   0  this   1    +1
550-600 ms  ┃                         main   1  this   0    -1
600-650 ms  ┃                         main   1  this   0    -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▂█▁····  139.1 (-33%)  130 (-32%)  242 (-26%)   396 (-9%)  3000
sweep    ······▂█▁▁···  145.7 (-49%)  126 (-47%)  454 (-35%)  680 (-54%)  3000
gw 1x    ·····▁▃█▁····  127.2 (-41%)  122 (-36%)  175 (-55%)  323 (-49%)  5295
eve 1x   ·····▁▂█▂▁▁▁·    223 (-23%)  133 (-32%)  271 (-57%)  3227 (+7%)  5186
eve 2x   ······▁█▃▁···  166.2 (-59%)  149 (-53%)  320 (-66%)  505 (-81%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▅▆▃▄▃▄█▁▄▁  126–158ms
sweep    ▄██▁▄▃▄▁▁█  127–170ms
gw 1x    ▃█▂▄▁▄▅▆▁▃  120–139ms
eve 1x   ▁▁▁▁▁▁▁▁█▄  133–714ms
eve 2x   ▃▂▂▁▂███▆▁  143–197ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  █▇▄▂▁▄▆  144–148ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▄▄▂▄▃▃█▃▃▁  37–51ms
sweep    ▁█▅▄▁▁▄▄█▃  49–68ms
gw 1x    ▄█▃▁▃▅▅▅▃▃  32–45ms
eve 1x   ▁▁▁▁▂▂▁▂█▃  26–70ms
eve 2x   █▆▃▃▅▇▄▁▃▇  20–33ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 42 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ step-vs-timer-early-settlement completed 8 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 265.9 KiB (±0) 97.2 KiB (-26 B) 1.99 MiB (+3.4 KiB)
nextjs-turbopack 273.3 KiB (±0) 426 B (±0) 981.6 KiB (+1.9 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

70520fb · run

@alangenfeld alangenfeld left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Client side of the capability flag. See the comment on vercel/workflow-server#1102: splitting replies unconditionally breaks existing clients on replies between 12 and 16 MiB. The client half is to send the capability on the upgrade request, and optionally to split requests only when the server confirms support.

Smaller items:

  • events-v4.ts: recordWsPartCounts was inserted between wsReplyStatus and its doc comment. The "Read the status off a reply frame…" comment now sits directly above this function's own comment instead of above wsReplyStatus.
  • wsMaxMessageBytes has no upper limit, so a value above 16 MiB brings back the silent 1006. It also silently ignores bad values, where the repo's envNumber clamps and warns once. Suggest envNumber with min: 1024, max: 16 MiB. #4510 reads the same variable for stream WebSocket writes that way, so whichever PR lands second should move both transports to one parser.
  • WORKFLOW_WS_MAX_MESSAGE_BYTES is documented in worlds/v5/vercel.mdx but not in docs/v5/configuration/worlds.mdx, which lists WORKFLOW_EVENTS_TRANSPORT.

Nits (optional):

  • decodeFrame fits better in frames.ts, where the stream session's async one-frame decoder could use it too.
  • request() holds about twice the payload in memory at peak: once from buildFrame, again from splitEncodedFrame.
  • The part assembler accepts a first part for any reqId, not only pending ones.
  • There's no test for a send error on a later part, or for the socket closing partway through a split frame.

Local agent review (`anthropic/claude-opus-5.5`)

@alangenfeld alangenfeld left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. With the server deploying ahead of the client, the documented rollout order covers split requests. A few follow-ups inline; none block.


Local agent review (`anthropic/claude-opus-5.5`)

Comment thread packages/world-vercel/src/ws-transport.ts
Comment thread packages/world-vercel/src/ws-parts.ts Outdated
Comment thread packages/world-vercel/src/ws-parts.ts

@shalabhc shalabhc left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. Addressed in 2b8f0e2, 62ea839 and f209e32:

  • Capability flag: every upgrade sends x-workflow-ws-flags: frame-parts, and the backend splits replies only for clients that list it. The header is a list, so later protocol flags can reuse it. Requests are always split. The upgrade response can't carry a confirmation (experimental_upgradeWebSocket writes the 101 itself), so the backend ships first and this release waits a day or two behind it; the backend PR has the rollout order.
  • Doc comment: recordWsPartCounts moved above the wsReplyStatus doc comment, so that comment sits on its function again.
  • Limit: wsMaxMessageBytes uses envNumber (integer, clamped to 2–16 MiB). The floor is 2 MiB, not 1 KiB, because of the part-size minimum below.
  • Docs: WORKFLOW_WS_MAX_MESSAGE_BYTES added to docs/v5/configuration/worlds.mdx.
  • decodeFrame: moved to frames.ts.
  • Assembler limits: now the same as the backend's:
    • at most one open split frame per connection, of at most 256 MiB;
    • every part except the last must carry at least 1 MiB, so a frame has at most 257 parts, and a larger partCount is rejected up front;
    • a whole reply for a reqId whose split reply is open is a protocol error;
    • a split reply for a request that has already settled is read through without buffering, then dropped with a log line.
  • Tests: added for:
    • a send error on a later part;
    • a close partway through a split reply;
    • the flag on the upgrade;
    • a split reply for a settled request;
    • reqId reuse;
    • each limit.

Still open: request_parts is only recorded once a reply arrives. I'll move it before the send and add span tests in a follow-up. Peak memory in request() (the encoded frame plus its split copy) is also left for later.

vercel Bot and others added 6 commits September 30, 2026 20:07
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
…ress review

Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
… least 1 MiB

Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
@shalabhc
shalabhc force-pushed the switch-ai-gateway-transport-from-http-to-websock-split-frames branch from a217f4a to 4000211 Compare September 30, 2026 20:07
…hout buffering them

Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Comment thread .changeset/ws-split-large-frames.md Outdated
vercel Bot and others added 2 commits September 30, 2026 21:29
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
@shalabhc
shalabhc merged commit 7fb0575 into main Sep 30, 2026
184 of 185 checks passed
@shalabhc
shalabhc deleted the switch-ai-gateway-transport-from-http-to-websock-split-frames branch September 30, 2026 22:47
@github-actions github-actions Bot mentioned this pull request Sep 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 7fb0575 (AI decision).

The WebSocket events transport this change extends does not exist on stable — git ls-tree origin/stable -- packages/world-vercel/src/ has no ws-transport.ts (nor ws-parts.ts, ws-transport*.test.ts), and docs/content/worlds/ is absent there too, so the change builds entirely on main-only code. It is also feature work regardless: it adds a new wire-protocol capability (x-workflow-ws-flags: frame-parts split frames requiring backend support), a new config knob (WORKFLOW_WS_MAX_MESSAGE_BYTES), new exports, and new telemetry attributes.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

7fb057520d97ef32a9e3b717efd51e594924f964

This branch was successfully deployed

1 active deployment
Preview – workflow-docs — 70520fb8 Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants