[core] Add a retention option to start() - #3787
Conversation
Adds `start({ retention })` for expressing a data-retention preference for
after a run completes. It is the typed spelling of a new reserved
`$retention` run attribute, which Worlds read at terminal cleanup.
- 'default' is not written at all, so it stays exactly equivalent to
omitting the option and spends none of the per-run attribute budget.
- 'none' and custom World-specific strings are seeded onto the run, with
the reserved-namespace opt-in set so the server accepts the `$` key.
- The value goes through the same validation as caller attributes, so an
oversized custom value fails in start() rather than at the World boundary.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 19f858d The changes in this PR will be included in the next version bump. This PR includes changesets to release 21 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 147324ms → this run 128373ms (Δ -18951ms, -13%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
🧪 E2E Test Results✅ All tests passed
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ✅ ▲ Vercel Production | 3662 | 0 | 685 | 4347 |
| ✅ 💻 Local Development | 3922 | 0 | 586 | 4508 |
| ✅ 📦 Local Production | 3922 | 0 | 586 | 4508 |
| ✅ 🐘 Local Postgres | 3922 | 0 | 586 | 4508 |
| ✅ 🪟 Windows | 320 | 0 | 2 | 322 |
| ✅ 🌐 Cross-language Conformance | 68 | 0 | 74 | 142 |
| ✅ vercel-http-transport | 823 | 0 | 143 | 966 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 557 | 0 | 87 | 644 |
| Total | 17223 | 0 | 2749 | 19972 |
Details by Category
✅ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 133 | 0 | 28 |
| ✅ astro-quickjs | 133 | 0 | 28 |
| ✅ example-node | 133 | 0 | 28 |
| ✅ example-quickjs | 133 | 0 | 28 |
| ✅ express-node | 133 | 0 | 28 |
| ✅ express-quickjs | 133 | 0 | 28 |
| ✅ fastify-node | 133 | 0 | 28 |
| ✅ fastify-quickjs | 133 | 0 | 28 |
| ✅ hono-node | 133 | 0 | 28 |
| ✅ hono-quickjs | 133 | 0 | 28 |
| ✅ nest-node | 133 | 0 | 28 |
| ✅ nest-quickjs | 133 | 0 | 28 |
| ✅ nextjs-turbopack-node | 158 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 158 | 0 | 3 |
| ✅ nextjs-webpack-node | 158 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 158 | 0 | 3 |
| ✅ nitro-node | 133 | 0 | 28 |
| ✅ nitro-quickjs | 133 | 0 | 28 |
| ✅ nuxt-node | 133 | 0 | 28 |
| ✅ nuxt-quickjs | 133 | 0 | 28 |
| ✅ python-node | 66 | 0 | 95 |
| ✅ sveltekit-node | 152 | 0 | 9 |
| ✅ sveltekit-quickjs | 152 | 0 | 9 |
| ✅ tanstack-start-node | 133 | 0 | 28 |
| ✅ tanstack-start-quickjs | 133 | 0 | 28 |
| ✅ vite-node | 133 | 0 | 28 |
| ✅ vite-quickjs | 133 | 0 | 28 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 27 |
| ✅ astro-stable-quickjs | 134 | 0 | 27 |
| ✅ express-stable-node | 134 | 0 | 27 |
| ✅ express-stable-quickjs | 134 | 0 | 27 |
| ✅ fastify-stable-node | 134 | 0 | 27 |
| ✅ fastify-stable-quickjs | 134 | 0 | 27 |
| ✅ hono-stable-node | 134 | 0 | 27 |
| ✅ hono-stable-quickjs | 134 | 0 | 27 |
| ✅ nest-stable-node | 134 | 0 | 27 |
| ✅ nest-stable-quickjs | 134 | 0 | 27 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 20 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 20 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 1 |
| ✅ nitro-stable-node | 134 | 0 | 27 |
| ✅ nitro-stable-quickjs | 134 | 0 | 27 |
| ✅ nuxt-stable-node | 134 | 0 | 27 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 27 |
| ✅ sveltekit-stable-node | 153 | 0 | 8 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 8 |
| ✅ tanstack-start-node | 134 | 0 | 27 |
| ✅ tanstack-start-quickjs | 134 | 0 | 27 |
| ✅ vite-stable-node | 134 | 0 | 27 |
| ✅ vite-stable-quickjs | 134 | 0 | 27 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 160 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 160 | 0 | 1 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 68 | 0 | 74 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 28 |
| ✅ express | 133 | 0 | 28 |
| ✅ hono | 133 | 0 | 28 |
| ✅ nextjs-turbopack | 158 | 0 | 3 |
| ✅ nitro | 133 | 0 | 28 |
| ✅ vite | 133 | 0 | 28 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 133 | 0 | 28 |
| ✅ express | 133 | 0 | 28 |
| ✅ nextjs-turbopack | 158 | 0 | 3 |
| ✅ vite | 133 | 0 | 28 |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
|
DX nits: would prefer will need to also enforce verification for retenion (for ex. what happens if the run is requesting larger retention than the team's plan supports? we can't reject the lazy start run without introducing server round trip for validation - so the run would have to be either failed, or the requested retention just gets overridden by server to |
The server reads `$retention` as a duration written as a decimal integer, not as the name of a mode: it honors the string `'0'`, and resolves everything else — `'none'` included — to the plan default while counting it as `retention.unsupported`. As written the two halves disagreed and the option would have been inert, so the wire value moves to `'0'`. The unit that duration is measured in has deliberately not been decided yet. It will most likely be seconds or milliseconds, chosen for granularity, and explicitly not days. Zero is the one value that means the same thing in every unit, which is exactly why it can ship ahead of that decision: it commits to a shape — a number, so the namespace has somewhere to grow — without committing to a scale. That is also why the option is typed `0 | 'default'` rather than `number | 'default'`. Someone writing `retention: 7` today has no unit to have meant it in, and the server would quietly keep their data; the literal type makes that a compile error, and a runtime guard makes it a thrown error for untyped JS callers. The arbitrary-string pass-through goes for the same reason — its only example, `'90d'`, baked in a unit — and callers targeting a World with its own retention vocabulary still have the documented escape hatch of writing `$retention` through `attributes` with `allowReservedAttributes`. With arbitrary strings gone, the empty-value and max-byte-length checks are unreachable (a fixed one-byte value cannot fail either), so they go too rather than sit as dead branches.
* origin/retention-postgres-world: [world-postgres] Honor $retention: 0 when a run finishes
* origin/retention-local-world: [world-local] Implement zero retention
Both World implementations arrived with their own copy of the $retention resolver, written independently against the same spec. They agreed today — same regex, same near-miss handling — but two hand-written parsers of a delete-or-keep predicate is a standing drift risk, and drift here means one World deleting a run another keeps. So the resolver moves to @workflow/world beside RETENTION_ATTRIBUTE, which already lived there, and both Worlds call it. world-local keeps its wrapper because it adds something real: a dev-server warning on an unrecognized value, which is where a developer can still notice an SDK sending a dialect this version predates. The parser's tests move with it, and gain coverage for the convenience predicate the Worlds actually call. They belong with the shared code, not with whichever World happened to be written first. Also adds the two missing changesets: world-postgres had none, so its implementation would not have shipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both Worlds also handle streams, and the Postgres purge is transactional — worth saying, since the whole page is about what guarantee you actually get. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The test pads its input past the inline-ref cutoff so the World writes a real blob to delete, which also crosses the JS client's compression threshold. The Python SDK cannot read zstd, so the run fails during input deserialization.
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Peter Wielander <mittgfu@gmail.com>
|
No backport to This commit adds a brand-new public API surface — the To override, re-run the Backport to stable workflow manually via |
Adds a
retentionoption tostart()so callers can express a data-retention preference for after a run completes.How it works
The option is the typed spelling of a new reserved run attribute,
$retention, exported from@workflow/worldasRETENTION_ATTRIBUTE.start()seeds it onto the run at creation, alongside the existing$rootRunId/$parentRunIdlineage keys, so it costs no extra write and rides both therun_createdevent and the resilient-start queue input. Worlds read it when a run reaches a terminal state.Decisions worth reviewing:
'default'writes nothing. The docstring says it is the same as omission, so writing$retention: 'default'would only spend one of the 64 per-run attribute slots to express what an absent key already means. It also keepsretention: 'default'from erroring on a World that predates attributes.$retentionattribute. A caller passing bothattributes: { '$retention': ... }(withallowReservedAttributes) andretention:gets the latter.retentionis the supported spelling; the raw attribute is the escape hatch. This is the opposite precedence from lineage, where caller attributes win, because lineage is inferred context andretentionis an explicit choice.0requires spec version 4 or later and throws otherwise, matching how initialattributesbehave. Silently dropping a data-deletion preference would be worse than failing.0 | 'default', notnumber | 'default'. A caller writingretention: 7today has no unit to have meant it in, and every World would resolve it to its own default — so the literal type makes it a compile error, and a runtime guard makes it a thrown error for untyped JS callers.attributes: { '$retention': ... }withallowReservedAttributes.Server side
The Vercel World implementation reads this attribute at terminal cleanup and deletes S3 ref payloads and stream content instead of tagging them for lifecycle expiration. That is vercel/workflow-server#858. Until it ships,
$retention: '0'is recorded on the run and otherwise inert, which is the safe failure mode: data keeps the plan's retention.Other Worlds ignore the attribute entirely today.
Test plan
New unit tests in
packages/core/src/runtime/start-retention.test.ts(9 cases): the0->'0'encoding and the reserved-namespace opt-in on both creation paths,'default'being byte-identical to omission, coexistence with lineage and caller attributes, precedence over a hand-written attribute, the reserved-key rejection still firing without the escape hatch, the spec-version gate (and'default'not tripping it), and the runtime guard rejecting a non-zero duration and a string from an untyped caller.start.test.tsandstart-lineage.test.tspass unchanged.Update (2026-08-27): wire value is
0, not'none'Amending this draft to match the server contract that landed in vercel/workflow-server#858. The body above has been corrected in place; this section records what changed and why. The original design — the
$namespace opt-in, the option beating a hand-written$retention, the reserved-key rejection without the escape hatch, and the spec-version-4 requirement — is unchanged.The server no longer honors
'none'.$retentionis read as a duration written as a decimal integer, and the only value implemented is the string'0'. Anything else —'none'included — resolves to the plan default and increments aretention.unsupported{reason:malformed}counter. Left as it was, this PR and the server would have disagreed and the feature would have been inert.The unit is deliberately not decided yet. It will most likely be seconds or milliseconds, chosen for granularity, and explicitly not days. Zero is the one value that means the same thing in every unit, which is exactly why it can ship ahead of that decision: it commits to a shape — a number, so the namespace has somewhere to grow — without committing to a scale. Nothing in the option, the docs or the wire format names a unit, and nothing should until that call is made.
So the public option takes a number, typed as the literal
0:start({ retention })$retentionattribute0'0''default'0 | 'default'rather thannumber | 'default'is load-bearing: it is what stops someone writingretention: 7, meaning seven of something, and silently getting the World's default. Untyped JS callers hit a runtime guard with the same message.Dropped from the original design:
'90d'test. That example baked in a unit, which is the one thing this field must not do. Worlds with their own retention vocabulary still have theattributes+allowReservedAttributesescape hatch, which this PR already documents as the raw spelling.'0'neither can fail, so they were dead branches rather than guards.Verified:
pnpm typecheckat the repo root (43/43 tasks),pnpm lint(no new diagnostics),@workflow/coreunit suite (2283 passed, 3 expected-fail, 1 skipped),@workflow/worldunit suite (165 passed).e2e coverage for the purge itself
The unit tests only prove the SDK sends
$retention: '0'. Addede2e > retention > retention: 0 purges the run payloads once the run finishesinpackages/core/e2e/e2e.test.ts, which proves the data is afterwards gone: startaddTenWorkflowwithretention: 0, assert the real payloads, then poll until the run and step payloads read back as<data expired>and the run'sexpiredAtis in the past. The run itself is asserted to survive — only user data goes, so it stays listable in observability.Gated on
WORKFLOW_VERCEL_ENV, not!isLocalDeployment(): only the Vercel World implements the terminal-cleanup pass, and the latter predicate is also true for the Postgres lane.Which server this runs against. The four Vercel e2e lanes set
VERCEL_WORKFLOW_SERVER_URLfrom a secret on PRs and leave it empty onmain(.github/workflows/tests.yml:482-487). That secret points ate2e.vercel-workflow.com— a Vercel custom environment on the workflow-server production project that tracks workflow-servermainHEAD, currently serving98c48763(#858). So this test is meaningful on the PR. Onmainthe empty value falls back tohttps://vercel-workflow.com(packages/world-vercel/src/utils.ts:260-261) — production, which trails the e2e environment by roughly 20 minutes and does not yet contain #858. Worth confirming production has picked it up before merging, or the firstmainrun of this test can fail spuriously. The same applies tochangeset-release/*PRs, which deliberately get the empty value.🤖 Generated with Claude Code