[core] Fix Run.returnValue retrying its accessor step for a terminal run - #4326
Conversation
…al run The accessor is a built-in step, so what it throws is subject to the step executor's retry policy. A terminal run is immutable: once the accessor has successfully read one, re-running the body re-reads the same record. Today `WorkflowRunFailedError` / `WorkflowRunCancelledError` carry no fatal marker, so the executor spends the whole retry budget before the failure reaches the caller — and then replaces it with its retry-exhaustion `FatalError` wrapper, so `WorkflowRunFailedError.is(err)` is false in the caller's workflow. Closes #4288. Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
🦋 Changeset detectedLatest commit: 3730f7e The changes in this PR will be included in the next version bump. This PR includes changesets to release 20 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ✅ ▲ Vercel Production | 3670 | 0 | 731 | 4401 |
| ✅ 💻 Local Development | 4014 | 0 | 550 | 4564 |
| ✅ 📦 Local Production | 4014 | 0 | 550 | 4564 |
| ✅ 🐘 Local Postgres | 4014 | 0 | 550 | 4564 |
| ✅ 🪟 Windows | 324 | 0 | 2 | 326 |
| ✅ 🌐 Cross-language Conformance | 68 | 0 | 76 | 144 |
| ✅ vercel-http-transport | 825 | 0 | 153 | 978 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 559 | 0 | 93 | 652 |
| Total | 17515 | 0 | 2705 | 20220 |
Details by Category
✅ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 133 | 0 | 30 |
| ✅ astro-quickjs | 133 | 0 | 30 |
| ✅ example-node | 133 | 0 | 30 |
| ✅ example-quickjs | 133 | 0 | 30 |
| ✅ express-node | 133 | 0 | 30 |
| ✅ express-quickjs | 133 | 0 | 30 |
| ✅ fastify-node | 133 | 0 | 30 |
| ✅ fastify-quickjs | 133 | 0 | 30 |
| ✅ hono-node | 133 | 0 | 30 |
| ✅ hono-quickjs | 133 | 0 | 30 |
| ✅ nest-node | 133 | 0 | 30 |
| ✅ nest-quickjs | 133 | 0 | 30 |
| ✅ nextjs-turbopack-node | 160 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 160 | 0 | 3 |
| ✅ nextjs-webpack-node | 160 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 160 | 0 | 3 |
| ✅ nitro-node | 133 | 0 | 30 |
| ✅ nitro-quickjs | 133 | 0 | 30 |
| ✅ nuxt-node | 133 | 0 | 30 |
| ✅ nuxt-quickjs | 133 | 0 | 30 |
| ✅ python-node | 66 | 0 | 97 |
| ✅ sveltekit-node | 152 | 0 | 11 |
| ✅ sveltekit-quickjs | 152 | 0 | 11 |
| ✅ tanstack-start-node | 133 | 0 | 30 |
| ✅ tanstack-start-quickjs | 133 | 0 | 30 |
| ✅ vite-node | 133 | 0 | 30 |
| ✅ vite-quickjs | 133 | 0 | 30 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 29 |
| ✅ astro-stable-quickjs | 134 | 0 | 29 |
| ✅ express-stable-node | 134 | 0 | 29 |
| ✅ express-stable-quickjs | 134 | 0 | 29 |
| ✅ fastify-stable-node | 134 | 0 | 29 |
| ✅ fastify-stable-quickjs | 134 | 0 | 29 |
| ✅ hono-stable-node | 134 | 0 | 29 |
| ✅ hono-stable-quickjs | 134 | 0 | 29 |
| ✅ nest-stable-node | 134 | 0 | 29 |
| ✅ nest-stable-quickjs | 134 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 162 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 162 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 162 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 162 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 162 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 162 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 29 |
| ✅ nitro-stable-quickjs | 134 | 0 | 29 |
| ✅ nuxt-stable-node | 134 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 29 |
| ✅ sveltekit-stable-node | 153 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 10 |
| ✅ tanstack-start-node | 134 | 0 | 29 |
| ✅ tanstack-start-quickjs | 134 | 0 | 29 |
| ✅ vite-stable-node | 134 | 0 | 29 |
| ✅ vite-stable-quickjs | 134 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 29 |
| ✅ astro-stable-quickjs | 134 | 0 | 29 |
| ✅ express-stable-node | 134 | 0 | 29 |
| ✅ express-stable-quickjs | 134 | 0 | 29 |
| ✅ fastify-stable-node | 134 | 0 | 29 |
| ✅ fastify-stable-quickjs | 134 | 0 | 29 |
| ✅ hono-stable-node | 134 | 0 | 29 |
| ✅ hono-stable-quickjs | 134 | 0 | 29 |
| ✅ nest-stable-node | 134 | 0 | 29 |
| ✅ nest-stable-quickjs | 134 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 162 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 162 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 162 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 162 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 162 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 162 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 29 |
| ✅ nitro-stable-quickjs | 134 | 0 | 29 |
| ✅ nuxt-stable-node | 134 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 29 |
| ✅ sveltekit-stable-node | 153 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 10 |
| ✅ tanstack-start-node | 134 | 0 | 29 |
| ✅ tanstack-start-quickjs | 134 | 0 | 29 |
| ✅ vite-stable-node | 134 | 0 | 29 |
| ✅ vite-stable-quickjs | 134 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 29 |
| ✅ astro-stable-quickjs | 134 | 0 | 29 |
| ✅ express-stable-node | 134 | 0 | 29 |
| ✅ express-stable-quickjs | 134 | 0 | 29 |
| ✅ fastify-stable-node | 134 | 0 | 29 |
| ✅ fastify-stable-quickjs | 134 | 0 | 29 |
| ✅ hono-stable-node | 134 | 0 | 29 |
| ✅ hono-stable-quickjs | 134 | 0 | 29 |
| ✅ nest-stable-node | 134 | 0 | 29 |
| ✅ nest-stable-quickjs | 134 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 162 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 162 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 162 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 162 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 162 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 162 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 29 |
| ✅ nitro-stable-quickjs | 134 | 0 | 29 |
| ✅ nuxt-stable-node | 134 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 29 |
| ✅ sveltekit-stable-node | 153 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 10 |
| ✅ tanstack-start-node | 134 | 0 | 29 |
| ✅ tanstack-start-quickjs | 134 | 0 | 29 |
| ✅ vite-stable-node | 134 | 0 | 29 |
| ✅ vite-stable-quickjs | 134 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 162 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 162 | 0 | 1 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 68 | 0 | 76 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 30 |
| ✅ express | 133 | 0 | 30 |
| ✅ hono | 133 | 0 | 30 |
| ✅ nextjs-turbopack | 160 | 0 | 3 |
| ✅ nitro | 133 | 0 | 30 |
| ✅ vite | 133 | 0 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 30 |
| ✅ express | 133 | 0 | 30 |
| ✅ nextjs-turbopack | 160 | 0 | 3 |
| ✅ vite | 133 | 0 | 30 |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 157647ms → this run 141541ms (Δ -16106ms, -10%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
`WorkflowRunFailedError` and `WorkflowRunCancelledError` are only thrown after a run has been read *successfully* and found in a terminal state. A terminal run is immutable, so re-running the read returns the same record and throws the same error. Neither carried a fatal marker, so `Run.returnValue` — a built-in step — spent its whole retry budget doing exactly that when a parent awaited a child that had already failed, and the error the caller finally caught was the executor's retry-exhaustion `FatalError` wrapper rather than the `WorkflowRunFailedError` the docs point it at. Mark both non-retryable. Errors from *failing to read* the run (transport blips, `WorkflowRunNotFoundError` during a resilient start) are a different case and stay retryable. Closes #4288. Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
|
one edge case i think is worth pinning: the first accessor step now correctly sees the terminal-run error as non-retryable, but does that property survive the step boundary? from what i can tell, "WorkflowRunFailedError" falls through the generic error serializer, which keeps the name/message/stack/cause but drops custom fields like "fatal". so after the workflow catches the accessor failure, "WorkflowRunFailedError.is(error)" is still true, but "FatalError.is(error)" looks like it would be false. that seems observable if the workflow passes that error into another step and that step rethrows it — we'd start retrying the same terminal error again. would it be worth adding a round-trip assertion to the existing regression that "FatalError.is(error)" is still true after "failureOf()" hydrates it? that should tell us whether this needs serializer support or just narrower docs. |
…run (#4326) Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
|
Backport PR opened against |
Surface `createHook({ experimental_force: true })` (#4193) on the hook
conflict, idempotency, and event-sourcing pages, and note on the child
workflows section that a failed or canceled child's `returnValue` throws
on the first attempt without retrying (#4326).
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
* [docs] Cross-link hook takeover and terminal child-run errors
Surface `createHook({ experimental_force: true })` (#4193) on the hook
conflict, idempotency, and event-sourcing pages, and note on the child
workflows section that a failed or canceled child's `returnValue` throws
on the first attempt without retrying (#4326).
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
* [docs] Address review: use experimental_force for superseding the owner
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
---------
Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Co-authored-by: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
Closes #4288.
Run.prototype.returnValueis a built-in step, so whatever it throws is classified by the step executor's retry policy. When the target run is already terminal the accessor's read succeeds and it throwsWorkflowRunFailedError(orWorkflowRunCancelledErrorfor a cancelled run). Neither carried a fatal marker, so the executor treated them as transient and retried — even though a terminal run is immutable, so every retry re-reads the same record and throws the same error.Impact
Driving the accessor through every attempt (the issue's repro runs only the first) shows two consequences:
error-level log lines per await;FatalErrorwhose message isStep "…" failed after 3 retries: ….WorkflowRunFailedError.is(err)isfalsethere, and the error is only reachable viaerr.cause— which contradicts what theWorkflowRunFailedErrordocs tell users to write. The same await outside a workflow (not a step) surfacesWorkflowRunFailedErrordirectly, so the two contexts disagree today.Fix
Mark
WorkflowRunFailedErrorandWorkflowRunCancelledErrornon-retryable (fatal = true, the own-property opt-inFatalError.is()already recognises). Both are thrown only after a run has been read successfully and found terminal, so there is no case where retrying them can produce a different answer. Errors from failing to read the run — transport blips,WorkflowRunNotFoundErrorduring a resilient start — are a different case and stay retryable; a test pins that.WorkflowRunCancelledErroris included because it is the same accessor reaching the same conclusion about the same immutable state; leaving it out would fix half of one code path.Nothing else keys off these two classes:
classifyRunErrordoesn't consultFatalError, so runerrorCodeis unchanged (USER_ERROReither way), and the only other construction site is theonRunFailedlifecycle payload, which is not thrown for retry classification.Tests
packages/errors/src/terminal-run-error.test.ts— pins the marker on both classes and its absence onWorkflowRunNotCompletedError/WorkflowRunNotFoundError.packages/core/src/runtime/run-return-value-terminal-retry.test.ts— the end-to-end repro: seeds a terminal run in a realworld-local, registers the actualRun.prototype.returnValuegetter as a step, runs it through the realexecuteStep, and asserts the step emitsstep_created, step_started, step_failedwith nostep_retrying, carrying an intactWorkflowRunFailedError. A third case pins that a completed run still resolves normally.Both failed on
mainbefore the fix (expected 'retry' to be 'failed').Out of scope, noted while here
runIdanderrorCodedo not survive a step boundary: the genericErrorreducer carries onlyname/message/stack/cause, so no SDK error class without a dedicated reducer keeps its extra fields. That is orthogonal to the retry decision and affects every such class, so this PR identifies the run through the error message instead of widening into the serializer.Docs Preview
WorkflowRunFailedErrorWorkflowRunCancelledError(Behind deployment protection; requires Vercel team access.)
🤖 Generated with Claude Code