Skip to content

workflow: resilient start in the local world - #283

Merged
fantix merged 2 commits into
mainfrom
workflow-resilient-start-local
Aug 13, 2026
Merged

fantix merged 2 commits into
mainfrom
workflow-resilient-start-local

Conversation

@fantix

@fantix fantix commented Aug 11, 2026

Copy link
Copy Markdown
Member

Second of three. Stacked on #282.

A run_started for a run with no row is a legal state, not an error one — see #282 for why start() does not order its two writes. This teaches LocalWorld to recover, mirroring @workflow/world-local's events-storage.ts.

What's here

  • Create the run from the event's eventData when the row is missing. The create is exclusive, so a concurrent run_created from start() keeps the row and we continue on top of it rather than clobbering a run that may already be running.
  • Write the run_created the log never got, drawn into an earlier event-id slot so replay reads it first.
  • Strip eventData from the stored run_started — it belongs to run_created, and leaving it would put the run's input in the log twice.
  • Make run_started idempotent for a run already in running: return the run without appending a second event. workflow: resilient start in the local world #283 stops reading the run before deciding whether to write this event, so it will be issued on every delivery rather than only the first — a log that grew an entry per replay would be unbounded. This is the contract the upstream runtime comment names explicitly.

Two pre-existing divergences this required closing

Neither is caused by the change; both would have made the recovery path silently lossy.

  • A run_started or run_failed for a run that does not exist was being persisted anyway, leaving the caller to trip over the missing run later with no context. Now rejected, as the Vercel world's 404 does — upstream rejects these "to match the postgres and vercel worlds".
  • Run rows were losing encryptionPublicKey. The model did not carry it, so extra="ignore" dropped it on read and every lifecycle transition rewrote the row without it. Harmless while nothing wrote one, but resilient start is exactly the path where the key would be lost for the rest of the run's life — so the field is now modelled and forwarded through all four transitions.

Test plan

17 new tests. Each assertion was mutation-checked: breaking the exclusive write, the event ordering, the eventData strip, the spec-version passthrough, or the key forwarding each fails at least one test. One of my own tests was initially passing vacuously (it wrote the racing row before the call, so the recovery path was never entered) — the mutation run is what caught that.

uv run poe qa green at this commit.

@vercel

vercel Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
vercel-py Ready Ready Preview Aug 12, 2026 2:11am

Request Review

@fantix
fantix force-pushed the workflow-resilient-start-local branch from 54c3579 to bff1a8d Compare August 11, 2026 22:40
@fantix
fantix force-pushed the workflow-resilient-start-local branch from bff1a8d to 10bfb3b Compare August 12, 2026 01:47
@fantix
fantix force-pushed the workflow-resilient-start-local branch from 10bfb3b to 0b48738 Compare August 12, 2026 02:11
Base automatically changed from workflow-resilient-start-models to main August 13, 2026 22:57
fantix added 2 commits August 13, 2026 18:57
Create the run from the data the consumer forwards on `run_started` when
`run_created` never landed, mirroring `world-local`'s `events-storage.ts`.
The write is exclusive so a concurrent `run_created` keeps the row, and
the log gets the `run_created` it never got in an earlier slot.

Make `run_started` idempotent for a run already running: the consumer is
about to issue it on every delivery, not just the first.

Also reject a `run_started` / `run_failed` for a run that does not exist
rather than persisting it, and stop erasing `encryptionPublicKey` on
every lifecycle transition.
`world-local`'s transport is plain JSON, which has nowhere to put
`bytes`, so it smuggles them through as a `{__type: "Uint8Array"}`
envelope — `jsonReplacer` on the way out, `jsonReviver` on the way back.
Our file store already speaks that dialect via `dumps_js` / `read_json`;
the queue did not.

Nothing noticed while no queue payload carried bytes. `runInput.input`
does, so a run bootstrapped from a TS-written `run_started` was created
and then rejected by its own model: `RunInput.input` is `Any` and kept
the envelope, `NonFinalWorkflowRun.input` wants `bytes | str`.

Decode on receive, before anything validates the payload — which is what
lets `RunInput.input` stay as loose as TS's `z.unknown()` rather than
growing a per-field unwrapper. Encode on send for the same reason the
file store does: the receive handler re-sends the payload it was handed,
so decoding alone turns the re-enqueue path into a `TypeError`.
@fantix
fantix force-pushed the workflow-resilient-start-local branch from d966b24 to be2ea67 Compare August 13, 2026 22:57
@fantix
fantix merged commit 42dd697 into main Aug 13, 2026
14 checks passed
@fantix
fantix deleted the workflow-resilient-start-local branch August 13, 2026 23:00

This branch was successfully deployed

1 active (outdated) and 1 inactive deployments
ci — be2ea673 Deployed Aug 13, 2026 by fantix via test (py3.10) #1135
Preview — 0b487385 Deployed Aug 12, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants