Repository navigation
workflow: resilient start in the local world - #283
Merged
Merged
Conversation
Contributor
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
fantix
force-pushed
the
workflow-resilient-start-local
branch
from
August 11, 2026 22:40
54c3579 to
bff1a8d
Compare
fantix
force-pushed
the
workflow-resilient-start-local
branch
from
August 12, 2026 01:47
bff1a8d to
10bfb3b
Compare
fantix
force-pushed
the
workflow-resilient-start-local
branch
from
August 12, 2026 02:11
10bfb3b to
0b48738
Compare
fantix
force-pushed
the
workflow-resilient-start-local
branch
from
August 13, 2026 14:58
0b48738 to
d966b24
Compare
elprans
approved these changes
Aug 13, 2026
Create the run from the data the consumer forwards on `run_started` when `run_created` never landed, mirroring `world-local`'s `events-storage.ts`. The write is exclusive so a concurrent `run_created` keeps the row, and the log gets the `run_created` it never got in an earlier slot. Make `run_started` idempotent for a run already running: the consumer is about to issue it on every delivery, not just the first. Also reject a `run_started` / `run_failed` for a run that does not exist rather than persisting it, and stop erasing `encryptionPublicKey` on every lifecycle transition.
`world-local`'s transport is plain JSON, which has nowhere to put
`bytes`, so it smuggles them through as a `{__type: "Uint8Array"}`
envelope — `jsonReplacer` on the way out, `jsonReviver` on the way back.
Our file store already speaks that dialect via `dumps_js` / `read_json`;
the queue did not.
Nothing noticed while no queue payload carried bytes. `runInput.input`
does, so a run bootstrapped from a TS-written `run_started` was created
and then rejected by its own model: `RunInput.input` is `Any` and kept
the envelope, `NonFinalWorkflowRun.input` wants `bytes | str`.
Decode on receive, before anything validates the payload — which is what
lets `RunInput.input` stay as loose as TS's `z.unknown()` rather than
growing a per-field unwrapper. Encode on send for the same reason the
file store does: the receive handler re-sends the payload it was handed,
so decoding alone turns the re-enqueue path into a `TypeError`.
fantix
force-pushed
the
workflow-resilient-start-local
branch
from
August 13, 2026 22:57
d966b24 to
be2ea67
Compare
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Second of three. Stacked on #282.
A
run_startedfor a run with no row is a legal state, not an error one — see #282 for whystart()does not order its two writes. This teachesLocalWorldto recover, mirroring@workflow/world-local'sevents-storage.ts.What's here
eventDatawhen the row is missing. The create is exclusive, so a concurrentrun_createdfromstart()keeps the row and we continue on top of it rather than clobbering a run that may already be running.run_createdthe log never got, drawn into an earlier event-id slot so replay reads it first.eventDatafrom the storedrun_started— it belongs torun_created, and leaving it would put the run's input in the log twice.run_startedidempotent for a run already inrunning: return the run without appending a second event. workflow: resilient start in the local world #283 stops reading the run before deciding whether to write this event, so it will be issued on every delivery rather than only the first — a log that grew an entry per replay would be unbounded. This is the contract the upstream runtime comment names explicitly.Two pre-existing divergences this required closing
Neither is caused by the change; both would have made the recovery path silently lossy.
run_startedorrun_failedfor a run that does not exist was being persisted anyway, leaving the caller to trip over the missing run later with no context. Now rejected, as the Vercel world's 404 does — upstream rejects these "to match the postgres and vercel worlds".encryptionPublicKey. The model did not carry it, soextra="ignore"dropped it on read and every lifecycle transition rewrote the row without it. Harmless while nothing wrote one, but resilient start is exactly the path where the key would be lost for the rest of the run's life — so the field is now modelled and forwarded through all four transitions.Test plan
17 new tests. Each assertion was mutation-checked: breaking the exclusive write, the event ordering, the
eventDatastrip, the spec-version passthrough, or the key forwarding each fails at least one test. One of my own tests was initially passing vacuously (it wrote the racing row before the call, so the recovery path was never entered) — the mutation run is what caught that.uv run poe qagreen at this commit.