Skip to content

A WorkflowWorldError timeout is an unhandled rejection and kills the process (exit 128) #3935

Description

@fgascon

Summary

When a request to the Workflow world API times out, the resulting WorkflowWorldError surfaces as an unhandled rejection and takes the Node process down with exit status 128 — killing unrelated work in the same invocation.

The slow request itself is a platform-side matter I'm pursuing through Vercel support separately. This issue is only about the client behaviour: a timeout reading run metadata should be a rejection the caller can catch, not a process exit.

⨯ unhandledRejection: Error [WorkflowWorldError]:
  GET /v2/runs/wrun_01M1HJ78KDW4VPYD53X8G77N5C?remoteRefBehavior=resolve timed out after 71779ms
  at ignore-listed frames {
    status: undefined,
    code: 'TIMEOUT',
    url: 'https://vercel-workflow.com/api/v2/runs/wrun_01M1HJ78KDW4VPYD53X8G77N5C?remoteRefBehavior=resolve',
    retryAfter: undefined,
    [cause]: [Error [TimeoutError]: The operation was aborted due to timeout] { code: 23 }
  }
Node.js process exited with exit status: 128.

What made this costly

The invocation that died was handling a GitHub webhook — work with nothing to do with the run being resolved. Two runs timed out in that one invocation (~72s each) before the process exited:

  • wrun_01M1HJ78KDW4VPYD53X8G77N5C
  • wrun_01M1HJ4F5YKY1BTTXFZJP80WQV

Observed 2026-09-02T17:24:11Z, region yul1, on a Vercel deployment.

Because the rejection is unhandled rather than surfaced to the caller, there is no seam to add a timeout policy, a retry, or a degraded path: the process is simply gone, and the symptom presents as an unexplained function failure rather than as "reading run metadata was slow".

Expected

  • WorkflowWorldError with code: 'TIMEOUT' is delivered to the awaiting caller so it can be caught.
  • A failure to resolve run metadata does not terminate the process.
  • Ideally the timeout is configurable, or at least documented — 72s is long enough that a caller inside a function with its own duration limit will usually lose the invocation before the request gives up.

Environment

  • workflow 4.8.5 (also reproduced with the app pinned to 5.0.0-beta.47)
  • Node 24, Next.js 16, deployed on Vercel

Note

I do not yet know whether the underlying slowness is caused by unusually large runs in our project — a separate SDK bug (vercel/eve#2891) generated several hundred failing runs there the same day. Either way, the unhandled-rejection behaviour looks wrong independently of what made the request slow.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions