[world-vercel] Attest the executor's spec version on run_started - #4366
Conversation
run_started repeats the spec version the caller of start() stamped, which is older when the run was started from another deployment. Send the version this SDK mints alongside it, so the backend can move such a run onto the event-id scheme this runtime reads before any event is numbered. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 03e6e67 The changes in this PR will be included in the next version bump. This PR includes changesets to release 21 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (62 failed)python-node (62 failed):
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ❌ ▲ Vercel Production | 3816 | 62 | 739 | 4617 |
| ✅ 💻 Local Development | 4238 | 0 | 550 | 4788 |
| ✅ 📦 Local Production | 4238 | 0 | 550 | 4788 |
| ✅ 🐘 Local Postgres | 4238 | 0 | 550 | 4788 |
| ✅ 🪟 Windows | 340 | 0 | 2 | 342 |
| ✅ vercel-http-transport | 873 | 0 | 153 | 1026 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 591 | 0 | 93 | 684 |
| Total | 18361 | 62 | 2637 | 21060 |
Details by Category
❌ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 141 | 0 | 30 |
| ✅ astro-quickjs | 141 | 0 | 30 |
| ✅ example-node | 141 | 0 | 30 |
| ✅ example-quickjs | 141 | 0 | 30 |
| ✅ express-node | 141 | 0 | 30 |
| ✅ express-quickjs | 141 | 0 | 30 |
| ✅ fastify-node | 141 | 0 | 30 |
| ✅ fastify-quickjs | 141 | 0 | 30 |
| ✅ hono-node | 141 | 0 | 30 |
| ✅ hono-quickjs | 141 | 0 | 30 |
| ✅ nest-node | 141 | 0 | 30 |
| ✅ nest-quickjs | 141 | 0 | 30 |
| ✅ nextjs-turbopack-node | 168 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 168 | 0 | 3 |
| ✅ nextjs-webpack-node | 168 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 168 | 0 | 3 |
| ✅ nitro-node | 141 | 0 | 30 |
| ✅ nitro-quickjs | 141 | 0 | 30 |
| ✅ nuxt-node | 141 | 0 | 30 |
| ✅ nuxt-quickjs | 141 | 0 | 30 |
| ❌ python-node | 4 | 62 | 105 |
| ✅ sveltekit-node | 160 | 0 | 11 |
| ✅ sveltekit-quickjs | 160 | 0 | 11 |
| ✅ tanstack-start-node | 141 | 0 | 30 |
| ✅ tanstack-start-quickjs | 141 | 0 | 30 |
| ✅ vite-node | 141 | 0 | 30 |
| ✅ vite-quickjs | 141 | 0 | 30 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 170 | 0 | 1 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ hono | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ nitro | 141 | 0 | 30 |
| ✅ vite | 141 | 0 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ vite | 141 | 0 | 30 |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 156783ms → this run 140582ms (Δ -16201ms, -10%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mped spec version Address review on #4327: - Raise the capability probe budget from 2s to 10s. healthCheck() returns on the first answer, so only a miss pays; a cold but healthy older-major target (e.g. stable, spec 3) is now read from the probe instead of being stamped with the spec-6 guess it would reject. - Be honest about the miss floor: spec 6 is the lowest a v5 runtime executes, and is not safe for older-major targets. Log a once-per-process warning on a miss, and TODO lowering the floor once #4366 and workflow-server#1044 ship. - Record workflow.run.spec_version / spec_version_source and the probe's latency or error on the start() span. - Cap `wf inspect` replay's probed spec version at the CLI World's. - Tests: cold spec-3 target, end-to-end probe timeout, 'latest' to another deployment and to self, retention error naming the target, compression off below spec 5, explicit-version attribution. - Docs and changeset updates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
run_started sent mintedSpecVersion() per request, which re-reads WORKFLOW_SEALED_LOG, while the runtime validated the value createWorld captured. Flipping the env in-process could make an executor attest a version it was never validated for, and the backend would raise runs to it. createWorld now reads it once and threads it to event writes (internal APIConfig.mintedSpecVersion); direct callers without it fall back to a fresh read. From the formal-model review on #4406 (E1). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
| export const CAPABILITY_ONLY_SPEC_VERSIONS: ReadonlySet<number> = new Set([ | ||
| SPEC_VERSION_SUPPORTS_CBOR_QUEUE_TRANSPORT, | ||
| SPEC_VERSION_SUPPORTS_ATTRIBUTES, | ||
| SPEC_VERSION_SUPPORTS_COMPRESSION, |
There was a problem hiding this comment.
Non-blocking, doc only. The definition says a capability-only version affects only "the runtime executing it". Compression (5) also affects anyone else who reads the run's payloads: a caller awaiting returnValue, a parent run, or an older dashboard/CLI. It can't be reached today, because only v5 executors attest their version and they're at 6 or above, so raising a run past 5 always crosses slot identity and needs a fresh log. Could the doc say that "capability-only" also means "no reader outside the executor is affected"? Then whoever classifies the next version checks that too.
Local agent review (`vercel-ai-gateway/anthropic/claude-opus-5.5`)
pranaygp
left a comment
There was a problem hiding this comment.
Approving. The attested version is now the World's declared specVersion, fixed when the World is created, and it is sent only on run_started. The capability/structural classification treats any unclassified version as structural, and the test that enumerates every SPEC_VERSION_SUPPORTS_* constant guards that. world 21 and world-vercel events 65 pass locally. The two failing CI jobs (Python Conformance and python-node) are also cancelled on every recent main run. Shipping this before workflow-server#1044 is safe: older servers ignore the unknown meta key.
|
Qualifying my approval note: this is safe to ship before vercel/workflow-server#1044 only while #1044 is not partly rolled out with |
|
No backport to The change depends entirely on main-only APIs: To override, re-run the Backport to stable workflow manually via |
Companion to #4327 for #4251.
Problem
A run's spec version is stamped by whoever calls
start(), and on world-vercel it decides the run's event-id scheme for life. #4327 makes a v5 caller stamp the target deployment's version. It cannot help when the caller is an older SDK that is already deployed: a v4 deployment callingstart(wf, [state], { deploymentId: 'latest' })after the project upgrades to v5 stamps spec 3. The backend then numbers the run's events by ULID, and the v5 runtime executing it fails every delivery (Event id is not slot-numbered) until the run ends withMAX_DELIVERIES_EXCEEDED. This is the handoff pattern the upgrading-workflows cookbook recommends, and the calling code runs on the old deployment, so no v4 release can fix it for runs already in flight.Change
run_startednow carriesexecutorSpecVersion: the spec version this World declared whencreateWorldran (mintedSpecVersion(), read once, so it followsWORKFLOW_SEALED_LOGat creation). It is read once so the attestation always matches what the runtime validated; the formal model on #4406 (E1) showed that a per-request env read could over-claim.specVersiononrun_startedis unchanged and still repeats the caller's stamp, which resilient start relies on.The world-vercel backend uses it to raise a run stamped lower to this version: before any other event exists when the raise changes the event-id scheme (for example a v4-stamped run moving onto slot ids), and at any point before the run finishes when it only enables newer capabilities. The field is a new optional frame-meta key: a backend that does not know it drops it, so this can ship before or after the backend change, and without it the backend keeps its current behavior.
Only world-vercel is affected at runtime. world-local and world-postgres mint their own event ids and are unchanged.
@workflow/worldalso gets the classification the backend's raise depends on, next to the constants:CAPABILITY_ONLY_SPEC_VERSIONS(3, 4, 5, 8),STRUCTURAL_SPEC_VERSIONS(2, 6, 7) andcrossesStructuralSpecVersion(from, to). A running run may only be raised across capability-only versions. The check is an allow-list, so an unclassified future version is structural by default, and a test fails when aSPEC_VERSION_SUPPORTS_*constant is added without a classification. vercel/workflow-server#1044 mirrors it inversion-utils.ts.Ordering. This has to ship with v5 GA for v4 → v5 handoffs to be repaired from day one. If it slips, the backend could fall back to the client's User-Agent (it already parses
@workflow/world-vercel/<version>, and "world-vercel ≥ 5 needs slots" is fixed), but that is only worth building if this misses GA.Tests
events.test.ts:run_startedsendsexecutorSpecVersionequal tomintedSpecVersion()next to the caller'sspecVersion, and follows the sealed-log kill switch.spec-version.test.ts: everySPEC_VERSION_SUPPORTS_*constant is in exactly one set; slot identity and the sealed log are structural; capability-only versions can be crossed; an unclassified future version cannot.pnpm --filter @workflow/world-vercel test: 753 passed. Typecheck clean.🤖 Generated with Claude Code