Skip to content

[core] Fix Run.returnValue retrying its accessor step for a terminal run - #4326

Merged
VaguelySerious merged 2 commits into
mainfrom
fix-returnvalue-retry
Sep 24, 2026
Merged

VaguelySerious merged 2 commits into
mainfrom
fix-returnvalue-retry

Conversation

@pranaygp

@pranaygp pranaygp commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Closes #4288.

Run.prototype.returnValue is a built-in step, so whatever it throws is classified by the step executor's retry policy. When the target run is already terminal the accessor's read succeeds and it throws WorkflowRunFailedError (or WorkflowRunCancelledError for a cancelled run). Neither carried a fatal marker, so the executor treated them as transient and retried — even though a terminal run is immutable, so every retry re-reads the same record and throws the same error.

Impact

Driving the accessor through every attempt (the issue's repro runs only the first) shows two consequences:

  • the accessor body runs 4 times instead of 1, and the remote failure reaches the caller ~3s late (plus queue redelivery overhead in a real World), with three error-level log lines per await;
  • once the budget is spent the executor wraps the error, so what the caller's workflow finally catches is a FatalError whose message is Step "…" failed after 3 retries: …. WorkflowRunFailedError.is(err) is false there, and the error is only reachable via err.cause — which contradicts what the WorkflowRunFailedError docs tell users to write. The same await outside a workflow (not a step) surfaces WorkflowRunFailedError directly, so the two contexts disagree today.

Fix

Mark WorkflowRunFailedError and WorkflowRunCancelledError non-retryable (fatal = true, the own-property opt-in FatalError.is() already recognises). Both are thrown only after a run has been read successfully and found terminal, so there is no case where retrying them can produce a different answer. Errors from failing to read the run — transport blips, WorkflowRunNotFoundError during a resilient start — are a different case and stay retryable; a test pins that.

WorkflowRunCancelledError is included because it is the same accessor reaching the same conclusion about the same immutable state; leaving it out would fix half of one code path.

Nothing else keys off these two classes: classifyRunError doesn't consult FatalError, so run errorCode is unchanged (USER_ERROR either way), and the only other construction site is the onRunFailed lifecycle payload, which is not thrown for retry classification.

Tests

  • packages/errors/src/terminal-run-error.test.ts — pins the marker on both classes and its absence on WorkflowRunNotCompletedError / WorkflowRunNotFoundError.
  • packages/core/src/runtime/run-return-value-terminal-retry.test.ts — the end-to-end repro: seeds a terminal run in a real world-local, registers the actual Run.prototype.returnValue getter as a step, runs it through the real executeStep, and asserts the step emits step_created, step_started, step_failed with no step_retrying, carrying an intact WorkflowRunFailedError. A third case pins that a completed run still resolves normally.

Both failed on main before the fix (expected 'retry' to be 'failed').

Out of scope, noted while here

runId and errorCode do not survive a step boundary: the generic Error reducer carries only name/message/stack/cause, so no SDK error class without a dedicated reducer keeps its extra fields. That is orthogonal to the retry decision and affects every such class, so this PR identifies the run through the error message instead of widening into the serializer.

Docs Preview

Page Preview
WorkflowRunFailedError https://workflow-docs-git-fix-returnvalue-retry.vercel.sh/docs/api-reference/workflow-errors/workflow-run-failed-error
WorkflowRunCancelledError https://workflow-docs-git-fix-returnvalue-retry.vercel.sh/docs/api-reference/workflow-errors/workflow-run-cancelled-error

(Behind deployment protection; requires Vercel team access.)

🤖 Generated with Claude Code

…al run

The accessor is a built-in step, so what it throws is subject to the step
executor's retry policy. A terminal run is immutable: once the accessor has
successfully read one, re-running the body re-reads the same record. Today
`WorkflowRunFailedError` / `WorkflowRunCancelledError` carry no fatal marker,
so the executor spends the whole retry budget before the failure reaches the
caller — and then replaces it with its retry-exhaustion `FatalError` wrapper,
so `WorkflowRunFailedError.is(err)` is false in the caller's workflow.

Closes #4288.

Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
@vercel

vercel Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
example-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-nestjs-workflow Building Building Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workflow-docs Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC
workflow-web Ready Ready Preview, v0 Sep 23, 2026 7:13pm UTC

@changeset-bot

changeset-bot Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 3730f7e

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
Name Type
@workflow/errors Patch
@workflow/core Patch
@workflow/builders Patch
@workflow/cli Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/world-local Patch
@workflow/world-postgres Patch
@workflow/world-vercel Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • onRunFailed receives the hydrated error with errorCode and cause (nextjs-webpack · local-dev / local / node / stable)
  • plainModuleDoneHook resumed via plain API route (o2flow shape) (nextjs-webpack · local-dev / local / node / stable)
  • promiseAllWorkflow via pages router (nextjs-webpack · local-dev / local / node / stable)
  • sleepingWorkflow via pages router (nextjs-webpack · local-dev / local / node / stable)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • cold-start-warmup · suite warmup (tanstack-start) · at 19:51:17Z · abandoned wrun_01M37X3SE1QFB93S7GFRKV6AX9 · (+1 more)
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 19:56:55Z · abandoned wrun_01M37XEYH21VMJ3785JJ276KCV

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3670 0 731 4401
✅ 💻 Local Development 4014 0 550 4564
✅ 📦 Local Production 4014 0 550 4564
✅ 🐘 Local Postgres 4014 0 550 4564
✅ 🪟 Windows 324 0 2 326
✅ 🌐 Cross-language Conformance 68 0 76 144
✅ vercel-http-transport 825 0 153 978
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 559 0 93 652
Total 17515 0 2705 20220
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 133 0 30
✅ astro-quickjs 133 0 30
✅ example-node 133 0 30
✅ example-quickjs 133 0 30
✅ express-node 133 0 30
✅ express-quickjs 133 0 30
✅ fastify-node 133 0 30
✅ fastify-quickjs 133 0 30
✅ hono-node 133 0 30
✅ hono-quickjs 133 0 30
✅ nest-node 133 0 30
✅ nest-quickjs 133 0 30
✅ nextjs-turbopack-node 160 0 3
✅ nextjs-turbopack-quickjs 160 0 3
✅ nextjs-webpack-node 160 0 3
✅ nextjs-webpack-quickjs 160 0 3
✅ nitro-node 133 0 30
✅ nitro-quickjs 133 0 30
✅ nuxt-node 133 0 30
✅ nuxt-quickjs 133 0 30
✅ python-node 66 0 97
✅ sveltekit-node 152 0 11
✅ sveltekit-quickjs 152 0 11
✅ tanstack-start-node 133 0 30
✅ tanstack-start-quickjs 133 0 30
✅ vite-node 133 0 30
✅ vite-quickjs 133 0 30

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 29
✅ astro-stable-quickjs 134 0 29
✅ express-stable-node 134 0 29
✅ express-stable-quickjs 134 0 29
✅ fastify-stable-node 134 0 29
✅ fastify-stable-quickjs 134 0 29
✅ hono-stable-node 134 0 29
✅ hono-stable-quickjs 134 0 29
✅ nest-stable-node 134 0 29
✅ nest-stable-quickjs 134 0 29
✅ nextjs-turbopack-canary-node 162 0 1
✅ nextjs-turbopack-canary-quickjs 162 0 1
✅ nextjs-turbopack-stable-node 162 0 1
✅ nextjs-turbopack-stable-quickjs 162 0 1
✅ nextjs-webpack-canary-node 162 0 1
✅ nextjs-webpack-canary-quickjs 162 0 1
✅ nextjs-webpack-stable-node 162 0 1
✅ nextjs-webpack-stable-quickjs 162 0 1
✅ nitro-stable-node 134 0 29
✅ nitro-stable-quickjs 134 0 29
✅ nuxt-stable-node 134 0 29
✅ nuxt-stable-quickjs 134 0 29
✅ sveltekit-stable-node 153 0 10
✅ sveltekit-stable-quickjs 153 0 10
✅ tanstack-start-node 134 0 29
✅ tanstack-start-quickjs 134 0 29
✅ vite-stable-node 134 0 29
✅ vite-stable-quickjs 134 0 29

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 29
✅ astro-stable-quickjs 134 0 29
✅ express-stable-node 134 0 29
✅ express-stable-quickjs 134 0 29
✅ fastify-stable-node 134 0 29
✅ fastify-stable-quickjs 134 0 29
✅ hono-stable-node 134 0 29
✅ hono-stable-quickjs 134 0 29
✅ nest-stable-node 134 0 29
✅ nest-stable-quickjs 134 0 29
✅ nextjs-turbopack-canary-node 162 0 1
✅ nextjs-turbopack-canary-quickjs 162 0 1
✅ nextjs-turbopack-stable-node 162 0 1
✅ nextjs-turbopack-stable-quickjs 162 0 1
✅ nextjs-webpack-canary-node 162 0 1
✅ nextjs-webpack-canary-quickjs 162 0 1
✅ nextjs-webpack-stable-node 162 0 1
✅ nextjs-webpack-stable-quickjs 162 0 1
✅ nitro-stable-node 134 0 29
✅ nitro-stable-quickjs 134 0 29
✅ nuxt-stable-node 134 0 29
✅ nuxt-stable-quickjs 134 0 29
✅ sveltekit-stable-node 153 0 10
✅ sveltekit-stable-quickjs 153 0 10
✅ tanstack-start-node 134 0 29
✅ tanstack-start-quickjs 134 0 29
✅ vite-stable-node 134 0 29
✅ vite-stable-quickjs 134 0 29

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 29
✅ astro-stable-quickjs 134 0 29
✅ express-stable-node 134 0 29
✅ express-stable-quickjs 134 0 29
✅ fastify-stable-node 134 0 29
✅ fastify-stable-quickjs 134 0 29
✅ hono-stable-node 134 0 29
✅ hono-stable-quickjs 134 0 29
✅ nest-stable-node 134 0 29
✅ nest-stable-quickjs 134 0 29
✅ nextjs-turbopack-canary-node 162 0 1
✅ nextjs-turbopack-canary-quickjs 162 0 1
✅ nextjs-turbopack-stable-node 162 0 1
✅ nextjs-turbopack-stable-quickjs 162 0 1
✅ nextjs-webpack-canary-node 162 0 1
✅ nextjs-webpack-canary-quickjs 162 0 1
✅ nextjs-webpack-stable-node 162 0 1
✅ nextjs-webpack-stable-quickjs 162 0 1
✅ nitro-stable-node 134 0 29
✅ nitro-stable-quickjs 134 0 29
✅ nuxt-stable-node 134 0 29
✅ nuxt-stable-quickjs 134 0 29
✅ sveltekit-stable-node 153 0 10
✅ sveltekit-stable-quickjs 153 0 10
✅ tanstack-start-node 134 0 29
✅ tanstack-start-quickjs 134 0 29
✅ vite-stable-node 134 0 29
✅ vite-stable-quickjs 134 0 29

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 162 0 1
✅ nextjs-turbopack-quickjs 162 0 1

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 76

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 133 0 30
✅ express 133 0 30
✅ hono 133 0 30
✅ nextjs-turbopack 160 0 3
✅ nitro 133 0 30
✅ vite 133 0 30

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 133 0 30
✅ express 133 0 30
✅ nextjs-turbopack 160 0 3
✅ vite 133 0 30

📋 View full workflow run

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 3730f7e · Wed, 23 Sep 2026 20:14:32 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 2391 (+72%) 🔻 2562 🔴 (+47%) 🔻 2583 🔴 (+45%) 🔻 2610 🔴 (+14%) 30
TTFS stream 208 (+28%) 🔻 2523 🔴 (+48%) 🔻 2563 🔴 (+47%) 🔻 2624 🔴 (+34%) 🔻 30
TTFS hook + stream 2711 (+330%) 🔻 2935 🔴 (+41%) 🔻 2991 🔴 (+32%) 🔻 3314 🔴 (+25%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 570 (+16%) 🔻 898 (+51%) 🔻 1142 (+76%) 🔻 3715 (+64%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 2299 (+26%) 🔻 3888 (+40%) 🔻 4061 (+41%) 🔻 13050 (+51%) 🔻 10
STSO 1020 steps (inline) 101 (-11%) 140 (+1.4%) 157 (+3.3%) 205 (-82%) 💚 1019
WO 1020 steps 141727 (-10%) 141727 (-10%) 141727 (-10%) 141727 (-10%) 1
CRTT first chunk (pooled) 76 (+23%) 🔻 122 (+20%) 🔻 143 (-26%) 💚 208 (-18%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 123 (+50%) 200 (+14%) 258 (-25%) 396 (-52%) 174 (+24%) 10
size sweep (100/s, 160B-12KB) 95.5 (+19%) 200 (-18%) 326 (-13%) 570 (-12%) 154 (-28%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 99 (+8%) 143 (-27%) 172 (-93%) 338 (-92%) 200 (-38%) 3
replay eve-gpt-5.6-sol-2000t (1x) 84 (-52%) 143 (-23%) 180 (-33%) 310 (-90%) 271 (-87%) 2
replay eve-gpt-5.6-sol-2000t (2x) 89 (-6%) 261 (-29%) 338 (-26%) 593 (-3%) 277 (-25%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 157647ms → this run 141541ms (Δ -16106ms, -10%)

  100-150 ms  ██████████████████████┃█  main 899  this 879   -20
  150-200 ms  ██┃                       main  69  this 127   +58
  200-250 ms  ┃                         main  13  this   6    -7
  250-300 ms  ┃                         main   7  this   1    -6
  300-350 ms  ┃                         main   2  this   2    +0
  350-400 ms  ┃                         main   2  this   1    -1
  400-450 ms  ┃                         main   1  this   3    +2
  500-550 ms  ┃                         main   1  this   0    -1
  550-600 ms  ┃                         main   5  this   0    -5
  700-750 ms  ┃                         main   2  this   0    -2
  750-800 ms  ┃                         main   2  this   0    -2
  800-850 ms  ┃                         main   2  this   0    -2
  850-900 ms  ┃                         main   1  this   0    -1
  900-950 ms  ┃                         main   1  this   0    -1
1050-1100 ms  ┃                         main   1  this   0    -1
1100-1150 ms  ┃                         main   1  this   0    -1
1150-1200 ms  ┃                         main   7  this   0    -7
1200-1250 ms  ┃                         main   1  this   0    -1
1300-1350 ms  ┃                         main   1  this   0    -1
1450-1500 ms  ┃                         main   1  this   0    -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▂█▃····     158 (+5%)  145 (+21%)  258 (-25%)  396 (-52%)  3000
sweep    ······▂█▃▁···    159 (-10%)  138 (-12%)  326 (-13%)  570 (-12%)  3000
gw 1x    ·····▁▃█▁····  124.7 (-53%)  120 (-10%)  172 (-93%)  338 (-92%)  5295
eve 1x   ·····▁▅█▁▁···  120.5 (-33%)  104 (-25%)  180 (-33%)  310 (-90%)  5186
eve 2x   ·····▁▂█▅▁···  187.9 (-13%)   151 (-7%)  338 (-26%)   593 (-3%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▇█▆▂▃▂▃▂▂▁  137–196ms
sweep    █▆▅▄▃▁▇█▆▃  136–175ms
gw 1x    █▄▃▄▃▃█▁▃▂  115–140ms
eve 1x   ▄▁▅▁▁█▆▆▃▄  105–142ms
eve 2x   ▂▁▃▁▁▂▃█▄▂  149–308ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▁▆█▇█▇▁  156–161ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▇█▃▃▁▃▄▃▂▁  42–73ms
sweep    ▃▁█▇▅▄▆▃▅▁  52–70ms
gw 1x    █▃▄▁▃▂▆▃▂▃  35–48ms
eve 1x   ▄▄▅▁▄█▅▇▃▆  21–27ms
eve 2x   █▄▁▁▃▄▄▂▄▃  25–36ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 42 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ step-vs-timer-early-settlement completed 8 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 251.1 KiB (±0) 93.8 KiB (+309 B) 1.90 MiB (-565 B)
nextjs-turbopack 258.5 KiB (±0) 426 B (±0) 903.1 KiB (-353 B)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

3730f7e · run

`WorkflowRunFailedError` and `WorkflowRunCancelledError` are only thrown after
a run has been read *successfully* and found in a terminal state. A terminal
run is immutable, so re-running the read returns the same record and throws the
same error. Neither carried a fatal marker, so `Run.returnValue` — a built-in
step — spent its whole retry budget doing exactly that when a parent awaited a
child that had already failed, and the error the caller finally caught was the
executor's retry-exhaustion `FatalError` wrapper rather than the
`WorkflowRunFailedError` the docs point it at.

Mark both non-retryable. Errors from *failing to read* the run (transport
blips, `WorkflowRunNotFoundError` during a resilient start) are a different
case and stay retryable.

Closes #4288.

Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@kvnloo

kvnloo commented Sep 23, 2026

Copy link
Copy Markdown

one edge case i think is worth pinning: the first accessor step now correctly sees the terminal-run error as non-retryable, but does that property survive the step boundary?

from what i can tell, "WorkflowRunFailedError" falls through the generic error serializer, which keeps the name/message/stack/cause but drops custom fields like "fatal".

so after the workflow catches the accessor failure, "WorkflowRunFailedError.is(error)" is still true, but "FatalError.is(error)" looks like it would be false.

that seems observable if the workflow passes that error into another step and that step rethrows it — we'd start retrying the same terminal error again.

would it be worth adding a round-trip assertion to the existing regression that "FatalError.is(error)" is still true after "failureOf()" hydrates it? that should tell us whether this needs serializer support or just narrower docs.

@VaguelySerious
VaguelySerious merged commit 2694663 into main Sep 24, 2026
498 of 509 checks passed
@VaguelySerious
VaguelySerious deleted the fix-returnvalue-retry branch September 24, 2026 00:08
github-actions Bot added a commit that referenced this pull request Sep 24, 2026
…run (#4326)

Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport PR opened against stable: #4346. Merge conflicts were resolved by AI — please review carefully. (backport job run)

pranaygp added a commit that referenced this pull request Sep 24, 2026
Surface `createHook({ experimental_force: true })` (#4193) on the hook
conflict, idempotency, and event-sourcing pages, and note on the child
workflows section that a failed or canceled child's `returnValue` throws
on the first attempt without retrying (#4326).

Co-Authored-By: Claude <noreply@anthropic.com>

Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
pranaygp added a commit that referenced this pull request Sep 24, 2026
* [docs] Cross-link hook takeover and terminal child-run errors

Surface `createHook({ experimental_force: true })` (#4193) on the hook
conflict, idempotency, and event-sourcing pages, and note on the child
workflows section that a failed or canceled child's `returnValue` throws
on the first attempt without retrying (#4326).

Co-Authored-By: Claude <noreply@anthropic.com>

Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>

* [docs] Address review: use experimental_force for superseding the owner

Co-Authored-By: Claude <noreply@anthropic.com>

Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>

---------

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Co-authored-by: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>

This branch was successfully deployed

18 active deployments
Preview – workflow-swc-playground — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workflow-docs — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-astro-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-vite-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – example-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workflow-tarballs — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-express-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-hono-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workflow-web — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Preview – workbench-python-workflow — 3730f7e7 Deployed Sep 23, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Run.returnValue retries its accessor step for an already-failed run

4 participants