Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
1888c75
[core] Add a retention option to start()
VaguelySerious Aug 25, 2026
15ce521
[core] Encode start({ retention }) as an integer duration
pranaygp Aug 28, 2026
2a74f94
[core] e2e: prove retention: 0 actually deletes the payloads
pranaygp Aug 28, 2026
3aaef98
[core] Rename the option to experimental_retention
pranaygp Aug 28, 2026
de4b5b6
[docs] Document retention, and that the World is what enforces it
pranaygp Aug 28, 2026
0423546
[core] Throw RunExpiredError from returnValue instead of a placeholder
pranaygp Aug 28, 2026
d07c668
[world-postgres] Honor $retention: 0 when a run finishes
pranaygp Aug 28, 2026
0d02eed
[world-local] Implement zero retention
pranaygp Aug 28, 2026
4ab20e9
Merge world-postgres zero-retention implementation
pranaygp Aug 28, 2026
adc598f
Merge world-local zero-retention implementation
pranaygp Aug 28, 2026
0f97eb7
[world] Share one retention parser across the Worlds
pranaygp Aug 28, 2026
efb3caa
[docs] Match the World table to the shipped implementations
pranaygp Aug 28, 2026
9ce172a
Merge branch 'main' into peter/start-retention-option
VaguelySerious Sep 8, 2026
a539e57
Consolidate the retention changesets into one
VaguelySerious Sep 8, 2026
5bd5658
Sort the retention exports in @workflow/world's index
VaguelySerious Sep 8, 2026
6696926
Skip the retention purge e2e on the python workbench
VaguelySerious Sep 8, 2026
2d1c848
Apply suggestion from @pranaygp
pranaygp Sep 8, 2026
19f858d
Apply batched suggestions from code review
VaguelySerious Sep 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .changeset/lucky-donkeys-repeat.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
'@workflow/core': minor
'@workflow/errors': minor
'@workflow/world': minor
'@workflow/world-local': minor
'@workflow/world-postgres': minor
---

Add an `experimental_retention` option to `start()`: `experimental_retention: 0` asks the World to delete the run's user data as soon as the run completes or fails, while keeping the run itself listable. Implemented on the Vercel, Postgres and Local Worlds. Reading a run whose data has expired now throws `RunExpiredError` instead of resolving to a placeholder.
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@ Learn more about [`WorkflowReadableStreamOptions`](/docs/api-reference/workflow-
* When you provide `deploymentId`, the argument types and return type become `unknown` because the workflow function's types may differ across deployments.
* `attributes` seeds plaintext run metadata as part of creation and requires a World implementing spec version 4 or later. Keys that start with `$` are reserved for framework and library code; framework-level callers can pass `allowReservedAttributes: true` to seed reserved keys, with the same semantics as the [`setAttributes`](/docs/api-reference/workflow/set-attributes) option of the same name.
* `region` pins the new run to a specific region on Worlds with a regional dimension. The [Vercel World](/worlds/vercel#explicit-region-selection) then serves the run's storage, queue dispatch, and streams from that region. When you omit `region`, the run is pinned to the region where it was created. Worlds without regions ignore the option.
* `experimental_retention` asks the World to delete the run's user data as soon as the run completes or fails, instead of keeping it for the World's default window. `0` requests immediate deletion; `'default'` is identical to omitting the option. These are the only two values accepted — the value is a duration and zero is the only one implemented, and its unit is not yet decided. Recorded as the reserved `$retention` attribute, so it needs a World implementing spec version 4 or later. Retention is enforced by the World, not the SDK: the first-party Worlds implement it and a World that does not keeps the data. Note that `await run.returnValue` on a run started with `experimental_retention: 0` usually throws [`RunExpiredError`](/docs/errors/run-expired) rather than resolving, because the deletion races the read. See [Data retention](/docs/observability/retention).

<Callout type="info">
If `start()` throws `'start' received an invalid workflow function. Ensure the Workflow SDK is configured correctly and the function includes a 'use workflow' directive.`, the compiler did not transform the passed function as a workflow. The two most common causes are a missing `"use workflow"` directive or missing framework integration. See [start-invalid-workflow-function](/docs/errors/start-invalid-workflow-function).
Expand Down
85 changes: 85 additions & 0 deletions docs/content/docs/v5/errors/run-expired.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
---
title: run-expired
description: A run's data passed its retention boundary, so its result can no longer be read.
type: troubleshooting
summary: Read a run's result before it expires, or return it through a channel you control.
prerequisites:
- /docs/foundations/workflows-and-steps
related:
- /docs/observability/retention
- /docs/api-reference/workflow-api/start
- /docs/foundations/hooks
---

## Error

```text
Run "wrun_..." completed, but its data expired at 2026-08-28T05:02:35.009Z
and is no longer readable.
```

Thrown as a `RunExpiredError` from `await run.returnValue`.

## Why this happens

A run's payloads — its input, output and error, and those of its steps — are
kept only for as long as the World's retention policy says. Its *metadata* —
id, status, timestamps — usually outlives them. So a run can be readable as a
record while its result is already gone.

Rather than hand back a placeholder that is indistinguishable from a value the
workflow genuinely returned, `returnValue` throws.

Two ways to reach it:

- **The run was started with `experimental_retention: 0`.** Its data is
deleted the moment it reaches a terminal state, and that deletion races your
own read of the result — and generally wins. On these runs, expect this
error rather than treating it as an edge case. See
[Data retention](/docs/observability/retention).
- **The run simply aged out.** It finished long enough ago that the World's
default retention window has passed.

## How to respond

`RunExpiredError` is terminal. Retrying will not bring the data back, so catch
Catch the error and use the run's metadata to decide what you want to do.

```typescript lineNumbers
import { getRun } from "workflow/api"
import { RunExpiredError } from "workflow/errors"

export async function readResult(runId: string) {
try {
return await getRun(runId).returnValue
} catch (error) {
if (RunExpiredError.is(error)) { // [!code highlight]
// `runStatus` is the run's terminal status when the World still has
// it, so you can tell a successful run whose result is gone from a
// failed one whose error is gone.
if (error.runStatus === "completed") {
// The run succeeded; its result is simply no longer stored.
}
return null
}
throw error
}
}
```

The error carries `runId`, `runStatus` and `expiredAt` when the World reports
them.

### If you need the result of a zero-retention run

Do not read it back off the run. Send it somewhere you control while the run
is still executing — a step that writes it to your own store. That is the intended pattern
for `experimental_retention: 0`: the point of the option is that the platform
does not keep your data, so the platform cannot also be where you fetch it
from afterwards.

## Related

If the run is gone entirely — metadata included — the World reports it as
missing and you get a `WorkflowRunNotFoundError` instead. That means the
record itself has been cleaned up, not just its payloads.
2 changes: 1 addition & 1 deletion docs/content/docs/v5/observability/meta.json
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
{
"title": "Observability",
"pages": ["tracing", "attributes"]
"pages": ["tracing", "attributes", "retention"]
}
93 changes: 93 additions & 0 deletions docs/content/docs/v5/observability/retention.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
---
title: Data retention
description: Control how long a run's data is kept after it finishes.
type: reference
summary: Control how long a run's data is kept after the run ends.
prerequisites:
- /docs/foundations/workflows-and-steps
related:
- /docs/observability
- /docs/api-reference/workflow-api/start
---

A finished run leaves data behind: the inputs and outputs of the workflow and
each of its steps, the payloads on its event log, and anything written to its
streams. How long that data is kept is decided by the World you are running
on, not by the SDK.

`experimental_retention` on [`start()`](/docs/api-reference/workflow-api/start)
lets a run ask for a specific retention period, rather than the World's default.

## Deleting a run's data as soon as it ends

{/* @skip-typecheck: abbreviated usage; processDocumentWorkflow is the reader's own workflow */}
```typescript lineNumbers
const run = await start(processDocumentWorkflow, [documentId], {
experimental_retention: 0, // [!code highlight]
})
```

`0` asks the World to delete the run's **user data** the moment the run
completes or fails, rather than keeping it for the World's default window.

Two values are accepted today:

| Value | Meaning |
| --- | --- |
| `0` | Delete user data as soon as the run reaches a terminal state. |
| `'default'` | Use the World's default. Identical to omitting the option. |

<Callout type="warn">
The option is prefixed `experimental_` because both its name and the set of
values it accepts are expected to change.
</Callout>

## What is deleted, and what is not

**Deleted:** the run's input, output and error; every step's input, output and
error; the payloads on the event log; and stream contents.

**Kept:** the run, step and event records themselves — their ids, timestamps,
status, step names, and any [attributes](/docs/observability/attributes) you
set. They are kept for the World's default period so the run stays visible in
the CLI and web UI. A purged run is still listed and still traceable; its
payloads simply read back as expired.

Inspecting a purged run shows it as expired rather than failing. The Workflow
CLI renders the run's own input, output and error as `<data expired>`:

```bash
workflow inspect runs wrun_...
```

Step, hook and event payloads read back empty. On the Vercel World they also
render as `<data expired>`; on Worlds that clear the stored value outright
they simply show as empty. Either way the data is gone — the difference is
only in how the absence is labelled.

<Callout type="warn">
**You cannot read the return value of a run started with
`experimental_retention: 0`.** The deletion races your own read of the
result and generally wins, so `await run.returnValue` throws
[`RunExpiredError`](/docs/errors/run-expired) instead of resolving.

This is a known limitation. If you need the result, send it somewhere you
control, e.g. a step that writes it to your own store, rather than reading it back off the run.
</Callout>

`RunExpiredError` is not specific to `experimental_retention: 0`. Any run read
after its retention window has passed throws it, and the error carries
`runId`, `runStatus` and `expiredAt` when the World still has them — so a
caller can tell a successful run whose result is gone from a failed one whose
error is gone. If the run's metadata is gone too, the World reports the run as
missing and you get `WorkflowRunNotFoundError` instead.

## Retention is implemented by the World

The SDK records your preference; it does not enforce it. `start()` writes the
value onto the run as the reserved `$retention` attribute, and the World
decides what to do when the run ends.
A World that does not implement retention
**keeps the data**. If you need certainty that a specific World deletes your data,
confirm it against that World's own documentation rather than the presence of this
option.
108 changes: 108 additions & 0 deletions packages/core/e2e/e2e.test.ts
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
import fs from 'node:fs';
import path from 'node:path';
import { setTimeout as sleep } from 'node:timers/promises';
Expand Down Expand Up @@ -76,6 +76,17 @@
// enough to flake without being any better at catching the regression.
const RACE_WINNER_MAX_DURATION_MS = 8_000;
const EVENT_POLL_PAGE_SIZE = 100;
/**
* What a purged payload looks like in `workflow inspect --json`.
*
* The server replaces expired payloads with a devalue stub that hydrates to
* `{ expiredAt: "<ISO>" }`; the CLI recognizes it with core's `isExpiredStub`
* and swaps in its `ExpiredDataRef` placeholder, whose `toJSON()` is this
* string. Asserting on it therefore exercises the same matcher the CLI and
* the web UI use, one layer up — the World returns payloads as raw devalue
* bytes, so there is nothing to run the predicate against down there.
*/
const EXPIRED_DATA_JSON = '<data expired>';

function expectElapsedAtLeast(
actualMs: number,
Expand Down Expand Up @@ -4793,4 +4804,101 @@
}
);
});

// ==========================================================================
// retention
// ==========================================================================

/**
* `start({ experimental_retention: 0 })` seeds `$retention: '0'`, which a
* World that implements retention honors at terminal cleanup by deleting
* the run's user payloads. The unit tests in `start-retention.test.ts`
* cover the SDK's half — that the attribute is encoded and sent. This
* covers the half only a real World can answer: that the data is
* afterwards actually gone.
*
* Gated on `WORKFLOW_VERCEL_ENV` — the same marker `setupWorld` uses to
* choose the Vercel world — rather than on `!isLocalDeployment()`, which is
* also true for the Postgres lane. The Local and Postgres Worlds implement
* retention too, but their coverage lives in their own package tests where
* the storage can be inspected directly.
*/
describe.skipIf(!process.env.WORKFLOW_VERCEL_ENV)('retention', () => {
test(
'experimental_retention: 0 purges the run payloads once the run finishes',
{ timeout: 240_000 },
async () => {
// Padded past the ~422-byte inline-ref cutoff on purpose. Below it a
// payload lives inside the database row and is scrubbed in place;
// above it the World writes a blob and has to delete the object. A
// small payload exercises only the first path, and this feature's
// whole claim is about the second. `metadataFromHelperWorkflow`
// echoes its label, so one big argument puts a blob behind the run's
// input, its output, and the step's on both sides.
const label = `retention-purge-${'x'.repeat(2048)}`;
const run = await start(
await e2e('metadataFromHelperWorkflow'),
[label],
{
experimental_retention: 0,
}
);

// The purge races the caller's own read of the result and generally
// wins, so `returnValue` resolves after the data is already gone.
// What it must NOT do is hand back the expired-data placeholder as
// though the workflow had returned it — that is indistinguishable
// from a real result. It throws instead, carrying whatever metadata
// outlived the payloads so a caller can still tell success from
// failure.
//
// Accepting either outcome would make this assertion worthless, so it
// insists on the throw. If the client ever starts winning the race
// this test fails loudly, which is the right way to find out.
await expect(run.returnValue).rejects.toMatchObject({
name: 'RunExpiredError',
runId: run.runId,
runStatus: 'completed',
});

// That same race is why nothing is asserted about the payloads
// *before* the purge: there is no reliable window in which to read
// them.
const afterPurge = await cliInspectJsonUntil(
`runs ${run.runId} --withData`,
(json) => json?.output === EXPIRED_DATA_JSON,
{ timeoutMs: 180_000, intervalMs: 5_000 }
);
expect(afterPurge).toMatchObject({
runId: run.runId,
input: EXPIRED_DATA_JSON,
output: EXPIRED_DATA_JSON,
// Only user data goes. The run itself survives on the World's
// default retention so it stays listable in observability.
status: 'completed',
});

// Step payloads go with it, not just the run's own input/output.
const steps = await cliInspectJsonUntil(
`steps --runId ${run.runId} --withData`,
(json) =>
Array.isArray(json) &&
json.length > 0 &&
json.every((step: any) => step.output === EXPIRED_DATA_JSON),
{ timeoutMs: 60_000, intervalMs: 5_000 }
);
expect(steps.length).toBeGreaterThan(0);
for (const step of steps) {
expect(step.output).toBe(EXPIRED_DATA_JSON);
}

// And the run carries the marker the CLI and web UI gate their
// "<data expired>" rendering on: an `expiredAt` in the past.
const world = await getWorld();
const persisted = await world.runs.get(run.runId);
expect(persisted.expiredAt).toBeInstanceOf(Date);
expect(persisted.expiredAt?.getTime()).toBeLessThanOrEqual(Date.now());
}
);
});
});
30 changes: 30 additions & 0 deletions packages/core/src/runtime/run.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
import {
RunExpiredError,
WorkflowRunCancelledError,
WorkflowRunFailedError,
WorkflowRunNotCompletedError,
Expand Down Expand Up @@ -396,6 +397,35 @@ export class Run<TResult> {

/** @internal */
async #resolveTerminalReturnValue(run: WorkflowRun): Promise<TResult> {
// Expiry is checked before the status branches, and deliberately so.
//
// Past its retention boundary a run's payloads are gone but its metadata
// usually is not, so `run.output` and `run.error` hydrate to an
// expired-data placeholder rather than to anything the caller asked for.
// Handing that back as if it were the return value — or wrapping it as
// the `cause` of a WorkflowRunFailedError — is worse than failing: it is
// indistinguishable from the workflow having genuinely returned it.
//
// A run started with `experimental_retention: 0` reaches this almost
// immediately (its purge races, and usually beats, the caller's own read
// of the result). An ordinary run reaches it whenever it is read after
// the World's default window. Both are the same condition and get the
// same terminal, non-retryable error, carrying whatever metadata
// survived so the caller can still tell success from failure.
//
// When even the metadata is gone the World reports the run as missing and
// the caller gets WorkflowRunNotFoundError from the read above instead —
// there is nothing left here to describe.
const expiredAt = run.expiredAt;
if (expiredAt != null && expiredAt <= new Date()) {
throw new RunExpiredError(
`Run "${this.runId}" ${run.status === 'completed' ? 'completed' : `is ${run.status}`}, but its data expired at ${expiredAt.toISOString()} and is no longer readable.`,
this.runId,
run.status,
expiredAt
);
}

if (run.status === 'completed') {
const encryptionKey = await this.#getEncryptionKey(run);
return await hydrateWorkflowReturnValue(
Expand Down
Loading
Loading