Skip to content

Run.returnValue retries its accessor step for an already-failed run #4288

Description

@fantix

Summary

Reading an already-failed run's returnValue as a built-in step schedules retries of the accessor, even though the target run is terminal and cannot produce a different result.

The target workflow is not rerun. The unnecessary retries are on the caller's returnValue accessor step, delaying delivery of the remote failure to the caller.

Expected / actual

Expected: once the accessor successfully reads a terminal failed run, propagate that failure without retrying the accessor. Transient errors fetching the run should still be retryable.

Actual: the accessor throws WorkflowRunFailedError, which is treated as a non-fatal step error. This happens even when its hydrated cause is a FatalError.

Standalone reproduction

Tested with Node.js v24.13.0 and the pinned packages below.

This is a runtime-level reproduction: it seeds a failed run in a real local World, registers the actual Run.prototype.returnValue getter, and invokes the actual step executor. It intentionally bypasses the framework/compiler/queue setup, and runs only the first attempt, which is enough to observe the erroneous retry decision. No deployment, server, credentials, or mocks are needed.

In an empty directory:

npm init -y
npm install --save-exact @workflow/core@5.0.0-beta.54 @workflow/world-local@5.0.0-beta.46 @workflow/world@5.0.0-beta.37 @workflow/errors@5.0.0-beta.22

Save this as repro.mjs, then run node repro.mjs. The assertions describe the current buggy behavior.

import assert from 'node:assert/strict';
import { mkdtemp, rm } from 'node:fs/promises';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { Run, getRun } from '@workflow/core/runtime/run';
import { dehydrateRunError, dehydrateStepArguments, hydrateStepError } from '@workflow/core/serialization';
import { FatalError } from '@workflow/errors';
import { SPEC_VERSION_CURRENT } from '@workflow/world';
import { createWorld } from '@workflow/world-local';

// Internal imports let this run without a framework, compiler, or deployment.
const runModule = import.meta.resolve('@workflow/core/runtime/run');
const { executeStep } = await import(new URL('./step-executor.js', runModule));
const { setWorld } = await import(new URL('./world.js', runModule));
const { registerStepFunction } = await import(new URL('../private.js', runModule));

const dataDir = await mkdtemp(join(tmpdir(), 'workflow-return-value-'));
const world = createWorld({ dataDir });
setWorld(world);
const emit = (runId, eventType, eventData, correlationId) =>
  world.events.create(runId, {
    eventType, eventData, correlationId, specVersion: SPEC_VERSION_CURRENT,
  });
async function createRun(workflowName) {
  const input = await dehydrateStepArguments([], 'run', undefined);
  const { run } = await emit(null, 'run_created', {
    deploymentId: 'dpl_repro', workflowName, input,
  });
  await emit(run.runId, 'run_started', {});
  return run.runId;
}

try {
  // Seed a permanently failed remote run; no target workflow is executed.
  const targetId = await createRun('target');
  await emit(targetId, 'run_failed', {
    error: await dehydrateRunError(new FatalError('target failed permanently'), targetId, undefined),
    errorCode: 'USER_ERROR',
  });
  await assert.rejects(getRun(targetId).returnValue, (error) => {
    console.log('Accessor error:', error.name);
    console.log('Accessor error is fatal:', FatalError.is(error));
    console.log('Original cause is fatal:', FatalError.is(error.cause));
    return error.name === 'WorkflowRunFailedError' && FatalError.is(error.cause);
  });

  // Execute the actual built-in getter body through the real step executor.
  // Binding supplies the Run receiver that the compiler normally serializes.
  const stepName = 'step//./repro//returnValue';
  registerStepFunction(stepName,
    Object.getOwnPropertyDescriptor(Run.prototype, 'returnValue').get.bind(getRun(targetId)));
  const callerId = await createRun('caller');
  const stepId = 'step_return_value';
  await emit(callerId, 'step_created', {
    stepName, input: await dehydrateStepArguments({ args: [] }, callerId, undefined),
  }, stepId);
  const result = await executeStep({
    world, workflowRunId: callerId, workflowName: 'caller',
    workflowStartedAt: Date.now(), stepId, stepName,
  });
  const { data: events } = await world.events.list({ runId: callerId });
  const retry = events.find((e) => e.eventType === 'step_retrying');
  console.log('Executor result:', result.type);
  console.log('Step events:', events.filter((e) => e.correlationId === stepId).map((e) => e.eventType).join(', '));
  if (retry) {
    const error = await hydrateStepError(retry.eventData.error, callerId, undefined);
    console.log('Retry error:', error.name);
    assert.equal(error.name, 'WorkflowRunFailedError');
  }
  assert.equal(result.type, 'retry');
  assert.ok(retry);
} finally {
  setWorld(undefined);
  await rm(dataDir, { recursive: true, force: true });
}

Output:

Accessor error: WorkflowRunFailedError
Accessor error is fatal: false
Original cause is fatal: true
Executor result: retry
Step events: step_created, step_started, step_retrying
Retry error: WorkflowRunFailedError

Likely cause

The same logic is present on main at 101d3472f6f150f0900713b4340c95d422ed673c:

The terminal status, rather than whether the original cause happened to be fatal, seems like the relevant distinction here: retrying a successful read of an immutable failed run cannot help.

I searched existing issues and PRs for WorkflowRunFailedError, returnValue + retry/retries/fatal, and related child-run failures, but did not find an existing report of this behavior.

Activity

  1. added a commit that references this issue on Sep 23, 2026
    3730f7e
  2. pranaygp commented on Sep 23, 2026

    @pranaygp
    Contributor

    (AI)

    Severity: S3

    No data loss and no hung runs — the remote failure does reach the caller, just late and re-typed — and there is a workaround, so this is S3 rather than S2. It does sit on a common path (a parent awaiting a child run), and the error-type change is the sharper edge of the two symptoms:

    • The accessor body runs 4 times instead of 1. On a plain error the retry delay is 1s, so the failure surfaces ~3s late (plus queue redelivery in a real World), with three error-level log lines per await.
    • Once the budget is spent the executor wraps the error, so what the caller's workflow catches is FatalError: Step "…" failed after 3 retries: …. WorkflowRunFailedError.is(err) is false there; the error is only reachable through err.cause. That contradicts the WorkflowRunFailedError docs, and it means the same await run.returnValue behaves differently inside a workflow (where it is a step) than outside one (where it is not). Reading err.cause is the workaround until the fix ships.

    WorkflowRunCancelledError has the same defect, from the same accessor reading the same kind of immutable state.

    Reproduction

    Your repro reproduces as written on the pinned betas under Node 22 as well as 24 — same four output lines, Executor result: retry.

    Driving it past the first attempt (looping executeStep until it stops returning retry) is what surfaces the second symptom:

    attempt 1: retry
    attempt 2: retry
    attempt 3: retry
    attempt 4: failed
    
    Step events: step_created, step_started, step_retrying, step_started, step_retrying,
                 step_started, step_retrying, step_started, step_failed
    Accessor body executions (step_started): 4
    Wall time to surface the remote failure: 3039ms
    
    What the caller workflow finally sees:
      name:                            FatalError
      WorkflowRunFailedError.is(err):  false
      err.cause?.name:                 WorkflowRunFailedError
    

    The same behaviour reproduces against main through the real executeStep over a real world-local, and is now pinned as a regression test.

    Related issues

    No duplicate, and no open issue covering this. The nearest neighbours touch the same accessor for unrelated reasons: #618 (closed — a returnValue poll that stalled on Vercel World) and #3570 (merged — resolving returnValue through a World long poll). Neither concerns the retry classification.

    Fix

    #4326 — marks WorkflowRunFailedError and WorkflowRunCancelledError non-retryable, exactly as you suggested: on the terminal status of the run that was read, not on whether the hydrated cause happens to be fatal. Both are thrown only after a successful read of a terminal run, so retrying them cannot produce a different answer. Errors from failing to read the run (transport blips, WorkflowRunNotFoundError during a resilient start) stay retryable, and a test pins that. CI is green.

    One thing the PR deliberately does not fix: runId and errorCode still do not survive a step boundary. The generic Error reducer carries only name/message/stack/cause, so no error class without a dedicated reducer keeps its extra fields. That is independent of the retry decision and affects every such class, so it is left for its own change. After the fix the caller gets the right error type and the original cause; the run id is in the message.

    To see if this fix solves your issue, could you try to install the pre-release in your package.json? Just replace the workflow dep like so:

    {
      "dependencies": {
        "workflow": "https://workflow-tarballs-awnl5osg8.labs.vercel.dev/workflow.tgz",
        "@workflow/ai": "https://workflow-tarballs-awnl5osg8.labs.vercel.dev/workflow-ai.tgz"
      }
    }

    The workflow tarball pulls the patched @workflow/errors transitively, so that one line is enough. For the standalone repro in this issue, which depends on @workflow/errors directly, point that dep at https://workflow-tarballs-awnl5osg8.labs.vercel.dev/workflow-errors.tgz too.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions