Repository navigation
Run.returnValue retries its accessor step for an already-failed run #4288
Description
Activity
- added a commit that references this issue
on Sep 23, 2026 (AI)
Severity: S3
No data loss and no hung runs — the remote failure does reach the caller, just late and re-typed — and there is a workaround, so this is S3 rather than S2. It does sit on a common path (a parent awaiting a child run), and the error-type change is the sharper edge of the two symptoms:
- The accessor body runs 4 times instead of 1. On a plain error the retry delay is 1s, so the failure surfaces ~3s late (plus queue redelivery in a real World), with three
error-level log lines per await. - Once the budget is spent the executor wraps the error, so what the caller's workflow catches is
FatalError: Step "…" failed after 3 retries: ….WorkflowRunFailedError.is(err)isfalsethere; the error is only reachable througherr.cause. That contradicts theWorkflowRunFailedErrordocs, and it means the sameawait run.returnValuebehaves differently inside a workflow (where it is a step) than outside one (where it is not). Readingerr.causeis the workaround until the fix ships.
WorkflowRunCancelledErrorhas the same defect, from the same accessor reading the same kind of immutable state.Reproduction
Your repro reproduces as written on the pinned betas under Node 22 as well as 24 — same four output lines,
Executor result: retry.Driving it past the first attempt (looping
executeStepuntil it stops returningretry) is what surfaces the second symptom:attempt 1: retry attempt 2: retry attempt 3: retry attempt 4: failed Step events: step_created, step_started, step_retrying, step_started, step_retrying, step_started, step_retrying, step_started, step_failed Accessor body executions (step_started): 4 Wall time to surface the remote failure: 3039ms What the caller workflow finally sees: name: FatalError WorkflowRunFailedError.is(err): false err.cause?.name: WorkflowRunFailedErrorThe same behaviour reproduces against
mainthrough the realexecuteStepover a realworld-local, and is now pinned as a regression test.Related issues
No duplicate, and no open issue covering this. The nearest neighbours touch the same accessor for unrelated reasons: #618 (closed — a
returnValuepoll that stalled on Vercel World) and #3570 (merged — resolvingreturnValuethrough a World long poll). Neither concerns the retry classification.Fix
#4326 — marks
WorkflowRunFailedErrorandWorkflowRunCancelledErrornon-retryable, exactly as you suggested: on the terminal status of the run that was read, not on whether the hydrated cause happens to be fatal. Both are thrown only after a successful read of a terminal run, so retrying them cannot produce a different answer. Errors from failing to read the run (transport blips,WorkflowRunNotFoundErrorduring a resilient start) stay retryable, and a test pins that. CI is green.One thing the PR deliberately does not fix:
runIdanderrorCodestill do not survive a step boundary. The genericErrorreducer carries onlyname/message/stack/cause, so no error class without a dedicated reducer keeps its extra fields. That is independent of the retry decision and affects every such class, so it is left for its own change. After the fix the caller gets the right error type and the originalcause; the run id is in the message.To see if this fix solves your issue, could you try to install the pre-release in your package.json? Just replace the
workflowdep like so:{ "dependencies": { "workflow": "https://workflow-tarballs-awnl5osg8.labs.vercel.dev/workflow.tgz", "@workflow/ai": "https://workflow-tarballs-awnl5osg8.labs.vercel.dev/workflow-ai.tgz" } }The
workflowtarball pulls the patched@workflow/errorstransitively, so that one line is enough. For the standalone repro in this issue, which depends on@workflow/errorsdirectly, point that dep athttps://workflow-tarballs-awnl5osg8.labs.vercel.dev/workflow-errors.tgztoo.- The accessor body runs 4 times instead of 1. On a plain error the retry delay is 1s, so the failure surfaces ~3s late (plus queue redelivery in a real World), with three
Summary
Reading an already-failed run's
returnValueas a built-in step schedules retries of the accessor, even though the target run is terminal and cannot produce a different result.The target workflow is not rerun. The unnecessary retries are on the caller's
returnValueaccessor step, delaying delivery of the remote failure to the caller.Expected / actual
Expected: once the accessor successfully reads a terminal
failedrun, propagate that failure without retrying the accessor. Transient errors fetching the run should still be retryable.Actual: the accessor throws
WorkflowRunFailedError, which is treated as a non-fatal step error. This happens even when its hydratedcauseis aFatalError.Standalone reproduction
Tested with Node.js v24.13.0 and the pinned packages below.
This is a runtime-level reproduction: it seeds a failed run in a real local World, registers the actual
Run.prototype.returnValuegetter, and invokes the actual step executor. It intentionally bypasses the framework/compiler/queue setup, and runs only the first attempt, which is enough to observe the erroneous retry decision. No deployment, server, credentials, or mocks are needed.In an empty directory:
Save this as
repro.mjs, then runnode repro.mjs. The assertions describe the current buggy behavior.Output:
Likely cause
The same logic is present on
mainat101d3472f6f150f0900713b4340c95d422ed673c:The terminal status, rather than whether the original cause happened to be fatal, seems like the relevant distinction here: retrying a successful read of an immutable failed run cannot help.
I searched existing issues and PRs for
WorkflowRunFailedError,returnValue+ retry/retries/fatal, and related child-run failures, but did not find an existing report of this behavior.