fix: stop stamping spec 8 into runs an older runtime executes - #4415
Conversation
🦋 Changeset detectedLatest commit: 97172e5 The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 131526ms → this run 154890ms (Δ +23364ms, +18%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
🧪 E2E Test Results✅ All tests passed
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ✅ ▲ Vercel Production | 3878 | 0 | 739 | 4617 |
| ✅ 💻 Local Development | 4238 | 0 | 550 | 4788 |
| ✅ 📦 Local Production | 4238 | 0 | 550 | 4788 |
| ✅ 🐘 Local Postgres | 4238 | 0 | 550 | 4788 |
| ✅ 🪟 Windows | 340 | 0 | 2 | 342 |
| ✅ 🌐 Cross-language Conformance | 68 | 0 | 84 | 152 |
| ✅ vercel-http-transport | 873 | 0 | 153 | 1026 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 591 | 0 | 93 | 684 |
| Total | 18491 | 0 | 2721 | 21212 |
Details by Category
✅ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 141 | 0 | 30 |
| ✅ astro-quickjs | 141 | 0 | 30 |
| ✅ example-node | 141 | 0 | 30 |
| ✅ example-quickjs | 141 | 0 | 30 |
| ✅ express-node | 141 | 0 | 30 |
| ✅ express-quickjs | 141 | 0 | 30 |
| ✅ fastify-node | 141 | 0 | 30 |
| ✅ fastify-quickjs | 141 | 0 | 30 |
| ✅ hono-node | 141 | 0 | 30 |
| ✅ hono-quickjs | 141 | 0 | 30 |
| ✅ nest-node | 141 | 0 | 30 |
| ✅ nest-quickjs | 141 | 0 | 30 |
| ✅ nextjs-turbopack-node | 168 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 168 | 0 | 3 |
| ✅ nextjs-webpack-node | 168 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 168 | 0 | 3 |
| ✅ nitro-node | 141 | 0 | 30 |
| ✅ nitro-quickjs | 141 | 0 | 30 |
| ✅ nuxt-node | 141 | 0 | 30 |
| ✅ nuxt-quickjs | 141 | 0 | 30 |
| ✅ python-node | 66 | 0 | 105 |
| ✅ sveltekit-node | 160 | 0 | 11 |
| ✅ sveltekit-quickjs | 160 | 0 | 11 |
| ✅ tanstack-start-node | 141 | 0 | 30 |
| ✅ tanstack-start-quickjs | 141 | 0 | 30 |
| ✅ vite-node | 141 | 0 | 30 |
| ✅ vite-quickjs | 141 | 0 | 30 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 170 | 0 | 1 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 68 | 0 | 84 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ hono | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ nitro | 141 | 0 | 30 |
| ✅ vite | 141 | 0 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ vite | 141 | 0 | 30 |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
The harness starts runs as the deployment under test (VERCEL_DEPLOYMENT_ID is the target), so start() treats them as same-deployment and stamps this SDK's spec version. Since #4193 that is 8, and the Python SDK rejects every run above 7 ('1 validation error for RunStartedEvent spec_version: Input should be less than or equal to 7'), leaving it pending. Both Python lanes have timed out on every run since. For apps that declare an e2e-conformance.json, probe the target once with the health check and stamp the lower of its answer and this SDK's version. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
resumeHook() stamped hook_received with SPEC_VERSION_CURRENT whatever the run's version. Since #4193 that is 8, and a runtime built against a lower version rejects the event when it replays the run. The Python SDK rejects the whole event log ('data.3.hook_received.specVersion: Input should be less than or equal to 7'), so every hook test on the Python lanes timed out even once runs were stamped for the target. Stamp the lower of the run's version and this SDK's. The payload is already encoded for the run's version. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
The health check reports the spec version a runtime writes, not the highest it reads. For the Python SDK those are 2 and 7, and stamping 2 disables everything gated at 4 or later: start() rejects initial attributes, and the resilient-start path changes. Four conformance tests that passed at 7 before #4193 failed. e2e-conformance.json already declares what an app's runtime supports, so it now carries maxSpecVersion (7 for the pinned Python SDK) and the harness stamps the lower of that and this SDK's version. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
A maxSpecVersion above what the app's runtime accepts leaves every run pending, which reads as a stalled queue. Say so where the warmup stall is reported, and name the exact constant and lockfile to keep it in step with. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
… version Run.cancel() and the DEPLOYMENT_MISMATCH run_failed write to a run from a process that is often not its executor, and stamped SPEC_VERSION_CURRENT like hook_received did. A runtime that validates the whole log (the Python SDK) rejects it. cancelRun() already used the run's version, so the two cancel paths disagreed. Move the rule into specVersionForRunWrite() and use it for all three writes. Run.cancel() now reads the run to learn its version. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
9e8bb00 to
176fbd4
Compare
VaguelySerious
left a comment
There was a problem hiding this comment.
AI review: no blocking issues
| 'use step'; | ||
| const world = await this.#lazyWorldPromise; | ||
| // The caller is often not the run's executor, so stamp the run's version. | ||
| const run = await world.runs.get(this.runId, { resolveData: 'none' }); |
There was a problem hiding this comment.
AI Review: Nit
This adds a runs.get round trip to every external Run.cancel(), before the write. It's correct, and cancelRun() in runs.ts already pays the same read, so both cancel paths are now consistent. Cancel isn't a hot path, so this is only worth a mention. I checked it end to end against world-local with a scratch test: on a run at SPEC_VERSION_CURRENT - 1, run_cancelled is stored at that version, the run ends cancelled, and cancelling a missing run still fails with a not-found error.
| * event log over a single such event. A run with no recorded version gets this | ||
| * SDK's, as before. | ||
| */ | ||
| export function specVersionForRunWrite( |
There was a problem hiding this comment.
AI Review: Nit
cancelRun(), wakeUpRun() and reenqueueRun() in runtime/runs.ts still stamp run.specVersion directly, uncapped. For a run newer than this SDK, they stamp a version this SDK can't produce, which this helper caps. Routing them through specVersionForRunWrite would give all external writes one rule. It's harmless for the reader, so a follow-up is fine.
…rsion They stamped run.specVersion directly, so a run started by a newer SDK got a version this one cannot write. Route them through specVersionForRunWrite(), which now takes the fallback for a run with no recorded version, so these keep treating such a run as legacy. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
TooTallNate
left a comment
There was a problem hiding this comment.
AI Review: Approved — no blocking issues found in 97172e5.
Reviewed the target-spec E2E wrapper, cross-runtime event stamping, and the latest cancelRun/wakeUpRun/reenqueueRun changes, including preservation of the legacy fallback. The earlier review feedback is addressed.
Validation: all 107 tests passed across run-spec-version, run-cancel, runs, resume-hook.durable, deployment-guard, and e2e/utils. Both Python CI lanes and the E2E Required Check also pass.
The remaining red Tarballs Preview Smoke Checks job was cancelled while waiting for the preview deployment; its endpoint-verification step never ran. That check does not provide evidence of a code regression in this diff.
|
No backport to This fixes a breakage that only exists on To override, re-run the Backport to stable workflow manually via |
Summary & Motivation
Both Python lanes (
E2E Vercel Prod Tests (python - node)andE2E Python Conformance) have timed out on every run since #4193 raised the spec version to 8, which failsE2E Required Checkon every PR. The deployed Python SDK writes spec 2 and reads at most 7, and two TypeScript writes stamp 8 into runs it executes:Runs. The harness starts runs as the deployment under test (
VERCEL_DEPLOYMENT_IDis the target), sostart()treats each as a same-deployment start and stamps this SDK's version. The deployment rejects every run and it stayspending:e2e-conformance.jsonalready declares what a non-JS app's runtime supports, so it now also carriesmaxSpecVersion: the highest spec version the runtime accepts (7 for the pinned Python SDK). The harness stamps runs with the lower of that and this SDK's version. The health check cannot supply this number, since it reports the version a runtime writes (2 for Python), and stamping 2 disables everything gated at 4 or later, such as initial attributes. fix(core): stamp cross-deployment runs with the target deployment's spec version #4327 fixes the same class for real cross-deployment starts but cannot reach the harness, which never looks cross-deployment.Events written to someone else's run.
resumeHook()stampedhook_receivedwithSPEC_VERSION_CURRENTregardless of the run, so the run's runtime rejected its whole event log on replay (data.3.hook_received.specVersion: Input should be less than or equal to 7). This is not specific to the harness: any 5.x caller resuming a hook on a run executed by an older or non-JS runtime breaks that run the same way.Run.cancel()(run_cancelled) and the deployment guard'sDEPLOYMENT_MISMATCHrun_failedhad the same over-stamp;cancelRun()already used the run's version, so the two cancel paths disagreed. All three now go throughspecVersionForRunWrite(), which stamps the lower of the run's version and this SDK's.Run.cancel()reads the run to learn its version.cancelRun(),wakeUpRun(), andreenqueueRun()inruns.tsalready used the run's version but uncapped, so a run started by a newer SDK got a version this one cannot write; they now use the same helper, keeping their legacy fallback for a run with no recorded version.Test Plan
specVersionForRunWrite()itself;hook_received(resume-hook.durable.test.ts),run_cancelled(run-cancel.test.ts), and the mismatchrun_failed(deployment-guard.test.ts) are stamped with an older run's version, andcancelRun/wakeUpRun/reenqueueRuncap a newer run's at this SDK's while still treating an unversioned run as legacy. Each fails without its fix. The full@workflow/coreunit suite passes.workbench-python-workflowpreview with CI's environment. Without the changes, warmup never picked up a run. With both,promiseAllWorkflow, the hook tests (hookWorkflow,hookWorkflow is not resumable via public webhook endpoint,hookCleanupTestWorkflow,hookWithSleepWorkflow), the threesetAttributes > start:tests, andresilient start: addTenWorkflow completes when run_created returns 500pass. Those last four passed at 7 before feat(hooks):createHook({ experimental_force: true })takes a held token over #4193.