fix(core): stamp cross-deployment runs with the target deployment's spec version - #4327
Conversation
`start({ deploymentId })` targeting another deployment stamped the run
with the caller World's spec version, even though the target executes
it. Resolve it from the capability probe `start()` already awaits,
capped at the caller World's version; on a probe miss fall back to the
lowest version the runtime serves rather than the caller's. Redirected
`recreateRunFromExisting` replays no longer pin the source run's
version.
Closes #4251
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
🦋 Changeset detectedLatest commit: 29197a1 The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 164957ms → this run 164332ms (Δ -625ms, 0%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (62 failed)python-node (62 failed):
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ❌ ▲ Vercel Production | 3816 | 62 | 739 | 4617 |
| ✅ 💻 Local Development | 4068 | 0 | 549 | 4617 |
| ✅ 📦 Local Production | 4238 | 0 | 550 | 4788 |
| ✅ 🐘 Local Postgres | 4238 | 0 | 550 | 4788 |
| ✅ 🪟 Windows | 340 | 0 | 2 | 342 |
| ✅ vercel-http-transport | 873 | 0 | 153 | 1026 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 591 | 0 | 93 | 684 |
| Total | 18191 | 62 | 2636 | 20889 |
Details by Category
❌ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 141 | 0 | 30 |
| ✅ astro-quickjs | 141 | 0 | 30 |
| ✅ example-node | 141 | 0 | 30 |
| ✅ example-quickjs | 141 | 0 | 30 |
| ✅ express-node | 141 | 0 | 30 |
| ✅ express-quickjs | 141 | 0 | 30 |
| ✅ fastify-node | 141 | 0 | 30 |
| ✅ fastify-quickjs | 141 | 0 | 30 |
| ✅ hono-node | 141 | 0 | 30 |
| ✅ hono-quickjs | 141 | 0 | 30 |
| ✅ nest-node | 141 | 0 | 30 |
| ✅ nest-quickjs | 141 | 0 | 30 |
| ✅ nextjs-turbopack-node | 168 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 168 | 0 | 3 |
| ✅ nextjs-webpack-node | 168 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 168 | 0 | 3 |
| ✅ nitro-node | 141 | 0 | 30 |
| ✅ nitro-quickjs | 141 | 0 | 30 |
| ✅ nuxt-node | 141 | 0 | 30 |
| ✅ nuxt-quickjs | 141 | 0 | 30 |
| ❌ python-node | 4 | 62 | 105 |
| ✅ sveltekit-node | 160 | 0 | 11 |
| ✅ sveltekit-quickjs | 160 | 0 | 11 |
| ✅ tanstack-start-node | 141 | 0 | 30 |
| ✅ tanstack-start-quickjs | 141 | 0 | 30 |
| ✅ vite-node | 141 | 0 | 30 |
| ✅ vite-quickjs | 141 | 0 | 30 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 170 | 0 | 1 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ hono | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ nitro | 141 | 0 | 30 |
| ✅ vite | 141 | 0 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ vite | 141 | 0 | 30 |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
|
one edge case i'm not following: on a probe timeout we fall back to spec 6, but the motivating case explicitly includes older cross-deployment targets like stable at spec 3. so if that older target is healthy but cold enough to miss the 2s probe, don't we recreate the original over-stamp? the caller can already mint lower-version runs — the pre-CBOR test demonstrates that — so i'm not sure the caller runtime's compatibility floor also needs to become the fallback for an unknown target. would it be worth adding a regression for a new starter targeting a spec-3 deployment whose probe times out, and asserting that the fallback can't exceed what the unknown target may support? |
…mped spec version Address review on #4327: - Raise the capability probe budget from 2s to 10s. healthCheck() returns on the first answer, so only a miss pays; a cold but healthy older-major target (e.g. stable, spec 3) is now read from the probe instead of being stamped with the spec-6 guess it would reject. - Be honest about the miss floor: spec 6 is the lowest a v5 runtime executes, and is not safe for older-major targets. Log a once-per-process warning on a miss, and TODO lowering the floor once #4366 and workflow-server#1044 ship. - Record workflow.run.spec_version / spec_version_source and the probe's latency or error on the start() span. - Cap `wf inspect` replay's probed spec version at the CLI World's. - Tests: cold spec-3 target, end-to-end probe timeout, 'latest' to another deployment and to self, retention error naming the target, compression off below spec 5, explicit-version attribution. - Docs and changeset updates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
|
@kvnloo good catch, you're right. A healthy
I've corrected the description and the |
|
@pranaygp, going through your review summary:
|
Apply the same `>= 1` integer check `start()` uses, so a JSON reply with `specVersion: 0` cannot stamp the replay with version 0. Also spell out, on the probe-miss floor, the conditions the formal model (#4406, C3) found for lowering it (#4401). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
alangenfeld
left a comment
There was a problem hiding this comment.
LGTM. One thing to check before merging: all 7 createHook({ experimental_force }) e2e tests failed on Vercel prod (quickjs), and "a run can take over its own earlier hook" timed out on local sveltekit. Neither failure shows up on main or #4366. The prod errors are UND_ERR_REQ_RETRY on the queue send, so this is probably flakiness, but the PR changes the specVersion that force-claim gates on. Worth a green re-run first.
Local agent review (`vercel-ai-gateway/anthropic/claude-opus-5.5`)
| * target a caller can reach attests its version (a published beta that does | ||
| * not would be stamped below slot identity and fail). Even then a floor | ||
| * below `SPEC_VERSION_SUPPORTS_ATTRIBUTES` would refuse `attributes` on every | ||
| * miss. Tracked in vercel/workflow#4401. |
There was a problem hiding this comment.
Non-blocking: lowering this floor to 3 depends on the run-spec-version-upgrade flag in vercel/workflow-server#1044 being on everywhere, not just on #1044 being merged. The flag is default-off there. Could this note (or #4401) say so, so nobody lowers the floor while the server still ignores executorSpecVersion?
Local agent review (`vercel-ai-gateway/anthropic/claude-opus-5.5`)
pranaygp
left a comment
There was a problem hiding this comment.
Thanks for the quick round. Everything from the earlier review is addressed. Two things are left before I'd approve:
- A repeated miss keeps the deployment on the 2 s budget (inline). This reopens kvnloo's case for a stable target that answers slowly.
- Bundle Size Report fails because of this PR. The flow bundle grew about 55 KB raw (+13.4 KiB gzip) for hono and nextjs-turbopack against the ff80d62 baseline, which is over the 51,200-byte gate. About half the added source is comments. Please either trim it or apply
allow-bundle-size-growthdeliberately.
Missing tests for the cache:
- the per-run public key is stripped from the cached entry;
- a cache hit falls back to
getEncryptionKeyForRun; - TTL expiry and eviction;
- the 2 s budget is actually applied (the answer-after-miss test only advances 1 s).
Worth a sentence in the description: on a cache hit the caller fetches the run's symmetric key, where a probe would have let it seal to the run's public key. That trades an extra lookup and wider key access for skipping one queue round trip.
CI, other failures:
- Python Conformance and python-node: also cancelled on every recent
mainrun. - Local Dev nextjs-turbopack / sveltekit (stable): most likely flakes (an HMR test, and a force-claim test on the same-deployment path that never probes).
- Vercel Prod nextjs-turbopack quickjs: a queues-proxy
UND_ERR_REQ_RETRY.
None of these look caused by this PR; they need a rerun.
Once (1) and (2) are done this looks good to merge ahead of #1044. The over-stamp is live on main (spec 8), and this PR fixes it whenever the probe answers.
| const answered = probe?.format !== undefined; | ||
| byDeployment.delete(key); | ||
| byDeployment.set(key, { | ||
| at: Date.now(), |
There was a problem hiding this comment.
Every repeated miss rewrites at, so while starts to this deployment keep arriving within the 10-minute TTL, the entry stays fresh and the deployment never gets the 10 s budget back. A late answer is dropped instead of cached. The result: a stable (spec 3) target that takes more than 2 s to answer when cold is stamped 6 on every start, and it rejects those runs. That's kvnloo's case again. Suggest keeping the original miss timestamp when the entry was already a miss (so it expires and the next probe gets 10 s again), and/or caching an answer that arrives after the budget.
|
https://github.com/vercel/vqs-server/pull/770 @alangenfeld @VaguelySerious I wonder if this PR is related to the failing CI tests? |
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Every miss rewrote the cache entry's timestamp, so while starts to a deployment kept arriving it never got the full probe budget back, and a target that answers only after the 2 s retry budget (a cold stable target) was stamped with the spec-6 fallback on every start. A miss now keeps a deployment on the short budget for 60 s from the first miss. Also test that a cached answer looks the run key up rather than reusing the probed one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
There was a problem hiding this comment.
Approving. The remaining blocker, a repeated probe miss pinning a deployment to the 2 s budget, is fixed in 29197a10f, along with a regression test and a check that a cache hit looks up the run key. We're accepting the flow-bundle growth (+55 KB raw) for this correctness fix, since the spec-8 over-stamp is live on main, so I added allow-bundle-size-growth. The Local Dev nextjs-webpack canary failure is the HMR suite, which doesn't go through start(). It needs a rerun (the bundle check also needs one to pick up the label).
Non-blocking follow-ups:
- tests for cache TTL expiry and eviction;
- a sentence in the description on the key-access trade-off: on a cache hit the caller fetches the run's symmetric key instead of sealing to its public key.
|
No backport to This is a genuine hang fix on To override, re-run the Backport to stable workflow manually via |
closes #4251
Problem
start({ deploymentId })aimed at another deployment stamped the run withopts.specVersion ?? world.specVersion. That is the caller's compiled-in version, and it is written torun_created, the queue message'srunInput, and the queuespecVersionoption. The run is executed by the target, so every>=gate keyed on the run's version (CBOR transport 3, attributes 4, compression 5, slot identity 6, sealed-lognoop7, and hook force-claim 8 from #4193) was reasoning about the wrong deployment:SPEC_VERSION_MAX_SUPPORTED,requiresNewerWorldrejectsrun_started, afterstart()has already returned. The run stayspendingindefinitely and the caller is never told. feat(hooks):createHook({ experimental_force: true })takes a held token over #4193 (now onmain) mints 8, so a new starter aimed at any 7/7 deployment hits this. It is also reachable againststable(spec 3) targets.start(self, …, { deploymentId: 'latest' })self-upgrade pattern.Fix (option 1 from the issue)
start()already awaits a health-check probe for every cross-deployment start, and the response carries the target's own mintedspecVersion. The probe's responder runs inside the target, so the value also honours the target'sWORKFLOW_SEALED_LOGswitch. Until nowstart()read the probe's key, core version and hook-resume version, but not its spec version. The newresolveCrossDeploymentSpecVersionuses it:world.specVersion. Unchanged, and no probe is sentopts.specVersionwf inspectreplay relies on itspecVersionmin(probed, world.specVersion)min(SPEC_VERSION_SUPPORTS_EVENT_SOURCING, world.specVersion)specVersion(so it is spec 3+)min(SPEC_VERSION_SUPPORTS_CBOR_QUEUE_TRANSPORT, world.specVersion)min(SPEC_VERSION_SUPPORTS_SLOT_IDENTITY, world.specVersion), plus a once-per-process warningRelated changes:
recreateRunFromExistingdefaults to the source run'sspecVersiononly when the replay stays on the source run's deployment. When a replay is redirected (the dashboard's Replay passes adeploymentId), it leaves the version unset sostart()resolves it from the probe. An explicitoptions.specVersionstill wins.healthCheck()returns as soon as the target answers, so only a miss pays for the budget. The probe used to be an optimisation whose fallback always worked. It now decides the run's version, and a healthy target that is only cold (for example astabletarget) must not fall through to a guess it would reject. 10 s matches the budgetwf inspectuses for the same probe. What a deployment answers is fixed for that deployment, so an answer is reused for later starts to it in the same process (per World, deployment and namespace; 10 min TTL, 256 entries). Those starts skip the probe, and the run's public key comes from the regular key lookup. After a miss, later probes to that deployment get 2 s. So only the first start to a target that never answers waits the full 10 s.start()span recordsworkflow.run.spec_versionandworkflow.run.spec_version_source(same-deployment/explicit/probe/probe-unversioned/probe-malformed/probe-miss/no-probe-channel). A cross-deployment start also recordsworkflow.capability_probe.cached, and a fresh probe recordsworkflow.capability_probe.latency_msorworkflow.capability_probe.error. A probe miss logs oneruntimeLogger.warnper process.wf inspectreplay caps its spec version at the CLI World's version, matchingstart(), whether it comes from the probe or from the source run.start.mdxbullet and a note in the v5upgrading-workflows.mdx(written as shipped with [world-vercel] Attest the executor's spec version on run_started #4366 / vercel/workflow-server#1044: the respawned run runs at the spec version of the deployment it lands on). Changeset:@workflow/coreand@workflow/clipatch. It calls out thatattributes/experimental_retentionaimed at a spec < 4 target now throw instead of creating a dead run, and that a redirected dashboard replay is capped at the dashboard's own version.Why this has the best forward/backward compatibility
start()already waits for that probe. Older targets work without changes, and so do older starters, which keep their current behaviour.run_createdand the arguments, and must not claim a log format it could not have produced. An earlier revision of this description claimed the fallback "errs low, never high". That is only true for same-major targets.stable/ pre-spec-3 target rejects the run, which is the same outcome asmain. Spec 3 would servestablebut break every v5 target. So this PR makes misses rare (10 s first budget), cheap to repeat (cache, 2 s after a miss) and visible (span attributes and a warning), and keeps 6. Once [world-vercel] Attest the executor's spec version on run_started #4366 and vercel/workflow-server#1044 are deployed, a v5 executor raises an under-stamped run onrun_started, so the miss floor can drop to 3 with a 2 s budget. That is tracked in Lower the cross-deployment probe-miss floor to spec 3 once executors raise under-stamped runs #4401.opts.specVersionstill overrides everything, so existing callers that pass a version keep working.Alternatives considered
run_startedsettle the version, treatingrun_created's as provisional (option 2 in the issue). As a general mechanism it is rejected because it changes every World and the workflow-server, and it is too late for choices the caller commits first (argument compression, CBOR vs JSON transport, attribute seeding). A narrow form of it is being done separately and complements this PR: [world-vercel] Attest the executor's spec version on run_started #4366 (the executor attestsexecutorSpecVersiononrun_started) and vercel/workflow-server#1044 (the backend raises a run to the executor's full version). Crossing a structural version (slot identity 6, the sealed log 7, or any unclassified future one) needs a fresh log; raises across capability-only versions (3, 4, 5, 8) apply any time before the run finishes. That is the only fix that reaches v4 callers already deployed, which no caller-side change can. The two follow different rules: this PR caps at the caller, and Enable lazyDiscovery in workbench #1044 raises the run to the executor's version after the creation round-trip (or at creation, on resilient start).executionContext(for examplehookForceClaimReaderVersion, reverted in0feb2fdonhook-force-claim). Rejected so there is one versioning mechanism. Every gate would need its own attestation, while fixing the stamp fixes all gates at once.wf inspect's replay did before this PR). Rejected because a newer target could make an older caller stamp a version it cannot write or read back.start()when the probe misses. Rejected because it would break starts that work today: same-major targets on a miss are served correctly by the spec-6 floor, and targets predating the health check never answer.SPEC_VERSION_SUPPORTS_SLOT_IDENTITYas well (suggested in triage). Rejected because a target that reports 3 (astable4.x deployment) cannot read a 6-stamped run. A reported value is trusted even when it is below the runtime floor. The floor applies only when we know nothing.SPEC_VERSION_MAX_SUPPORTEDinstead ofworld.specVersion. Rejected, narrowly. The two are equal today. Capping at the World's minted version also respects the caller's ownWORKFLOW_SEALED_LOGkill switch, which is the more conservative choice.healthCheck()already returns on the first answer, so a longer single wait costs the same on a hit, and on a miss it does not discard a late answer to the first probe or send a second queue message.mainnow mints 8, so the over-stamp against 7/7 targets is live today and an answered probe fixes it. The per-deployment cache keeps the 10 s budget off the steady state in the meantime, and Lower the cross-deployment probe-miss floor to spec 3 once executors raise under-stamped runs #4401 tracks the switch.Not in this PR
noop, hook force-claim). Those gates readrun.specVersion, so these tests pin the stamp they read. The force-claim gate (feat(hooks):createHook({ experimental_force: true })takes a held token over #4193) is now onmain; its reader test lives with it. The new tests cover new starter → older executor, older starter → newer executor, pre-CBOR and pre-JSON targets, probe misses, explicit override, and target-named attribute errors.Tests
packages/core/src/runtime/start-cross-deployment-spec-version.test.ts(new, 27 tests). Since the last round: an 8/8 caller and a 7/7 target stamp 7 (the case live onmain), a malformed JSON reply stamps 3 with CBOR transport, an answer is reused for a second start to the same deployment, an answer after a miss is cached, and the second miss waits only the short budget. Besides the original cases, it covers a coldstable(spec 3) target answering after 5 s, which is stamped 3 (this test fails with the old 2 s budget), and an end-to-end probe timeout (stamped 6 in all three places, warns once across two starts, recordsprobe-missand the error). It also coversdeploymentId: 'latest'resolving to another deployment and to the caller, retention errors naming the target, a probe below 5 disabling compression even when the core version says gzip, span attribution of an explicit version, and the resolver'ssourcefor every branch.packages/core/src/runtime/runs.test.ts: three newrecreateRunFromExistingcases (same-deployment pin, redirected replay left unset, explicit override).main(spec 8):packages/coresrc/2703 passed, the CLIinspecttests 120 passed,pnpm --filter @workflow/core typecheck, andbiome checkare clean. The only warnings are two pre-existing complexity warnings.Triage and repro: https://taskmaster.playground-vercel.tools/s/ses_2ruvkbsezfsj
🤖 Generated with Claude Code