fix(core): repay a force-claim victim's wake independent of the claimer's log tail - #4398
Conversation
🦋 Changeset detectedLatest commit: a866bbe The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (62 failed)python-node (62 failed):
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ❌ ▲ Vercel Production | 3816 | 62 | 739 | 4617 |
| ✅ 💻 Local Development | 4238 | 0 | 550 | 4788 |
| ✅ 📦 Local Production | 4238 | 0 | 550 | 4788 |
| ✅ 🐘 Local Postgres | 4238 | 0 | 550 | 4788 |
| ✅ 🪟 Windows | 340 | 0 | 2 | 342 |
| ✅ vercel-http-transport | 873 | 0 | 153 | 1026 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| ✅ vercel-ws-transport | 591 | 0 | 93 | 684 |
| Total | 18361 | 62 | 2637 | 21060 |
Details by Category
❌ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 141 | 0 | 30 |
| ✅ astro-quickjs | 141 | 0 | 30 |
| ✅ example-node | 141 | 0 | 30 |
| ✅ example-quickjs | 141 | 0 | 30 |
| ✅ express-node | 141 | 0 | 30 |
| ✅ express-quickjs | 141 | 0 | 30 |
| ✅ fastify-node | 141 | 0 | 30 |
| ✅ fastify-quickjs | 141 | 0 | 30 |
| ✅ hono-node | 141 | 0 | 30 |
| ✅ hono-quickjs | 141 | 0 | 30 |
| ✅ nest-node | 141 | 0 | 30 |
| ✅ nest-quickjs | 141 | 0 | 30 |
| ✅ nextjs-turbopack-node | 168 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 168 | 0 | 3 |
| ✅ nextjs-webpack-node | 168 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 168 | 0 | 3 |
| ✅ nitro-node | 141 | 0 | 30 |
| ✅ nitro-quickjs | 141 | 0 | 30 |
| ✅ nuxt-node | 141 | 0 | 30 |
| ✅ nuxt-quickjs | 141 | 0 | 30 |
| ❌ python-node | 4 | 62 | 105 |
| ✅ sveltekit-node | 160 | 0 | 11 |
| ✅ sveltekit-quickjs | 160 | 0 | 11 |
| ✅ tanstack-start-node | 141 | 0 | 30 |
| ✅ tanstack-start-quickjs | 141 | 0 | 30 |
| ✅ vite-node | 141 | 0 | 30 |
| ✅ vite-quickjs | 141 | 0 | 30 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 142 | 0 | 29 |
| ✅ astro-stable-quickjs | 142 | 0 | 29 |
| ✅ express-stable-node | 142 | 0 | 29 |
| ✅ express-stable-quickjs | 142 | 0 | 29 |
| ✅ fastify-stable-node | 142 | 0 | 29 |
| ✅ fastify-stable-quickjs | 142 | 0 | 29 |
| ✅ hono-stable-node | 142 | 0 | 29 |
| ✅ hono-stable-quickjs | 142 | 0 | 29 |
| ✅ nest-stable-node | 142 | 0 | 29 |
| ✅ nest-stable-quickjs | 142 | 0 | 29 |
| ✅ nextjs-turbopack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-stable-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-canary-quickjs | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-node | 170 | 0 | 1 |
| ✅ nextjs-webpack-stable-quickjs | 170 | 0 | 1 |
| ✅ nitro-stable-node | 142 | 0 | 29 |
| ✅ nitro-stable-quickjs | 142 | 0 | 29 |
| ✅ nuxt-stable-node | 142 | 0 | 29 |
| ✅ nuxt-stable-quickjs | 142 | 0 | 29 |
| ✅ sveltekit-stable-node | 161 | 0 | 10 |
| ✅ sveltekit-stable-quickjs | 161 | 0 | 10 |
| ✅ tanstack-start-node | 142 | 0 | 29 |
| ✅ tanstack-start-quickjs | 142 | 0 | 29 |
| ✅ vite-stable-node | 142 | 0 | 29 |
| ✅ vite-stable-quickjs | 142 | 0 | 29 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 170 | 0 | 1 |
| ✅ nextjs-turbopack-quickjs | 170 | 0 | 1 |
✅ vercel-http-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ hono | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ nitro | 141 | 0 | 30 |
| ✅ vite | 141 | 0 | 30 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
✅ vercel-ws-transport
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ example | 141 | 0 | 30 |
| ✅ express | 141 | 0 | 30 |
| ✅ nextjs-turbopack | 168 | 0 | 3 |
| ✅ vite | 141 | 0 | 30 |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
…er's log tail A replay republished the victim's wake only while the forced hook_created was the last row the claimer itself wrote, so any own row landing between the creation and the wake (a step or wait terminal from another invocation, or, since #4392, a row the same suspension writes alongside the creation) hid the debt, and a victim parked only on `await hook` never woke. Every replay now republishes for each forced hook_created in the loaded log whose createdAt is within 24 hours (the queue's message retention and idempotency window), under the existing `hook-force-claim-<hookId>` key. The node:vm handler sends it alongside the suspension's writes and at most once per hook per invocation; QuickJS once per invocation on load. Both use the shared rule in hook-wake.ts. Closes #4393. Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
Both engines used to create forced hooks one token at a time, each group waiting on its victim wake before the next started. That barrier is gone (#4392); these tests hold the first forced create until the second is issued, so a serial implementation deadlocks, and check that each victim is still woken once under its own key. Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>
840df4f to
2db67e3
Compare
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 149219ms → this run 147656ms (Δ -1563ms, -1%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
|
No backport to This fixes the victim-wake recovery for the To override, re-run the Backport to stable workflow manually via |
Follow-up to #4392 (merged); rebased onto
main. Closes #4393.Server-side model: vercel/workflow-server#1059.
Problem
A run that force-claims a token writes a forced
hook_createdand then publishes a wake for the run it took the token from. If the invocation dies between the two, the only recovery wasforcedCreationOwingWake: on the next replay, republish if the forced creation is still the last row the claimer wrote itself. Any own row landing after the creation hid the debt:A victim parked only on
await hookthen never woke.Fix (option 1 from the issue)
forcedCreationsOwingWake(events, nowMs)inhook-wake.tsnow returns every forcedhook_created(witheventData.forceClaimedFrom) in the loaded log whosecreatedAtis withinFORCE_CLAIM_WAKE_REPUBLISH_WINDOW_MS(24 h).republishOwedForceClaimVictimWakesrepublishes each one under the existinghook-force-claim-<hookId>key. It never looks at what came after the creation, so no row can hide the debt. Both engines share this rule.What is unchanged:
workflowNameare still skipped. On the replay path they are now skipped silently, because the creating invocation already logged the warning.hook_receivedand a sealed-lognoopnever mattered to the window rule.What changed in the engines:
handleSuspension): the republish no longer blocks the suspension. The rule does not depend on write order, so it runs alongside the suspension's writes and is joined before return (try/finallyaroundsettlePhase). A per-invocationforceClaimVictimWakesset, created inruntime.tsand also filled bycreateHookEventfor the invocation's own forced creations, keeps it to one send per hook per invocation across the replay loop's suspensions.The window bound: 24 h from the row's
createdAtretentionSecondsdefaults to 86400; world-vercel does not override it), and the creation is written after the send. So every such redelivery arrives within 24 h of the creation. That covers VQS's own redelivery backoff too (≤900 s per hop,MAX_QUEUE_DELIVERIES= 48, about 9-10 h).min(retention, 24h)(the@vercel/queueSendOptions.idempotencyKeydocs), so every republish inside the window collapses into the one delivered wake. world-postgres's handler skips a key it has already completed (an in-process LRU of 10k entries). world-local dedupes a key only while its message is in flight, so there a republish can deliver the victim one extra replay, which reads nothing new.event.createdAt, never the event id: a slot id decodes to epoch 0. On slot-identity runscreatedAtis the writer's clientoccurredAt, which workflow-server clamps to within 1 h ahead / 7 d behind server time (slotEventCreatedAt). Normal skew is milliseconds to seconds; worst case the window's edge moves by up to 1 h. A future-dated creation counts as recent, and an unparseablecreatedAtcounts as recent too (errs toward a wake). A republish that slips past VQS's window is one extra wake, which is harmless.Why not options 2 or 3
The window closes the gap with neither, and its only imprecision is a harmless duplicate wake.
Docs, tests, changeset
create-hook.mdx: the force-claim paragraph now describes a durable wake, republished within 24 h under one idempotency key. It no longer says "repaired on a best-effort basis". Doc comments inhook-wake.ts,suspension-handler.tsandquickjs-entrypoint.tsno longer point at Force-claim victim wake recovery shouldn't depend on the claimer's log tail #4393.suspension-handler.test.ts('force-claim victim wake'):step_created/step_completed/wait_completedhook_receivedhook_disposed{forceClaimedBy}, then an own row)hook_disposedquickjs-force-claim-wake.test.ts: recent creation,hook_received, own step/wait rows, chain, outside the window, and the same concurrent-forced-tokens test.@workflow/corepatch.Assumptions / open points
handleSuspension, so a replay that finishes the run without suspending does not republish. The same was true of the tail rule. QuickJS republishes on load and is not affected.Docs Preview
createHook(): taking over a token another run holdsChecks
cd packages/core && pnpm vitest run src: 2673 passedtsc --noEmitfor core: okbiome checkon the changed files: no new diagnostics🤖 Generated with Claude Code