Skip to content

fix(core): stamp cross-deployment runs with the target deployment's spec version - #4327

Merged
VaguelySerious merged 7 commits into
mainfrom
fix-cross-deploy-spec-version
Sep 25, 2026
Merged

VaguelySerious merged 7 commits into
mainfrom
fix-cross-deploy-spec-version

Conversation

@VaguelySerious

@VaguelySerious VaguelySerious commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

closes #4251

Problem

start({ deploymentId }) aimed at another deployment stamped the run with opts.specVersion ?? world.specVersion. That is the caller's compiled-in version, and it is written to run_created, the queue message's runInput, and the queue specVersion option. The run is executed by the target, so every >= gate keyed on the run's version (CBOR transport 3, attributes 4, compression 5, slot identity 6, sealed-log noop 7, and hook force-claim 8 from #4193) was reasoning about the wrong deployment:

  • New starter → older executor: the run is over-stamped. Once the gap passes the executor's SPEC_VERSION_MAX_SUPPORTED, requiresNewerWorld rejects run_started, after start() has already returned. The run stays pending indefinitely and the caller is never told. feat(hooks): createHook({ experimental_force: true }) takes a held token over #4193 (now on main) mints 8, so a new starter aimed at any 7/7 deployment hits this. It is also reachable against stable (spec 3) targets.
  • Older starter → newer executor: the run is under-stamped. Version-gated behaviour stays off for that run's whole life. This is the documented start(self, …, { deploymentId: 'latest' }) self-upgrade pattern.

Fix (option 1 from the issue)

start() already awaits a health-check probe for every cross-deployment start, and the response carries the target's own minted specVersion. The probe's responder runs inside the target, so the value also honours the target's WORKFLOW_SEALED_LOG switch. Until now start() read the probe's key, core version and hook-resume version, but not its spec version. The new resolveCrossDeploymentSpecVersion uses it:

Situation Stamped version
same deployment world.specVersion. Unchanged, and no probe is sent
explicit opts.specVersion that value. Unchanged; the CLI's wf inspect replay relies on it
probe reports specVersion min(probed, world.specVersion)
target answered in the pre-JSON plain-text format (pre-spec-3) min(SPEC_VERSION_SUPPORTS_EVENT_SOURCING, world.specVersion)
target answered in JSON without a usable specVersion (so it is spec 3+) min(SPEC_VERSION_SUPPORTS_CBOR_QUEUE_TRANSPORT, world.specVersion)
probe timed out (10 s for the first probe to a deployment, 2 s after a miss), or the World has no probe channel min(SPEC_VERSION_SUPPORTS_SLOT_IDENTITY, world.specVersion), plus a once-per-process warning

Related changes:

  • Attribute / retention gates now check the resolved version. When a probe lowers it, the error names the target deployment and the version it runs, rather than claiming the caller's World is too old.
  • recreateRunFromExisting defaults to the source run's specVersion only when the replay stays on the source run's deployment. When a replay is redirected (the dashboard's Replay passes a deploymentId), it leaves the version unset so start() resolves it from the probe. An explicit options.specVersion still wins.
  • Probe budget: 10 s for the first probe to a deployment, with a per-deployment cache. healthCheck() returns as soon as the target answers, so only a miss pays for the budget. The probe used to be an optimisation whose fallback always worked. It now decides the run's version, and a healthy target that is only cold (for example a stable target) must not fall through to a guess it would reject. 10 s matches the budget wf inspect uses for the same probe. What a deployment answers is fixed for that deployment, so an answer is reused for later starts to it in the same process (per World, deployment and namespace; 10 min TTL, 256 entries). Those starts skip the probe, and the run's public key comes from the regular key lookup. After a miss, later probes to that deployment get 2 s. So only the first start to a target that never answers waits the full 10 s.
  • Observability. The start() span records workflow.run.spec_version and workflow.run.spec_version_source (same-deployment / explicit / probe / probe-unversioned / probe-malformed / probe-miss / no-probe-channel). A cross-deployment start also records workflow.capability_probe.cached, and a fresh probe records workflow.capability_probe.latency_ms or workflow.capability_probe.error. A probe miss logs one runtimeLogger.warn per process.
  • wf inspect replay caps its spec version at the CLI World's version, matching start(), whether it comes from the probe or from the source run.
  • Docs: start.mdx bullet and a note in the v5 upgrading-workflows.mdx (written as shipped with [world-vercel] Attest the executor's spec version on run_started #4366 / vercel/workflow-server#1044: the respawned run runs at the spec version of the deployment it lands on). Changeset: @workflow/core and @workflow/cli patch. It calls out that attributes / experimental_retention aimed at a spec < 4 target now throw instead of creating a dead run, and that a redirected dashboard replay is capped at the dashboard's own version.

Why this has the best forward/backward compatibility

  • No protocol, schema or World change. It uses a probe field that every target since [cli] [core] Probe deployment specVersion before CLI start #1629 already sends, and start() already waits for that probe. Older targets work without changes, and so do older starters, which keep their current behaviour.
  • When the target answers, it is exact. The result is the target's own version, capped at the caller's World version. The caller writes run_created and the arguments, and must not claim a log format it could not have produced. An earlier revision of this description claimed the fallback "errs low, never high". That is only true for same-major targets.
  • A probe miss is a guess, and no single guess is safe. Spec 6 is the lowest version a v5 runtime can execute on the Vercel World, because it requires slot event ids. So spec 6 keeps v5 targets working, but a stable / pre-spec-3 target rejects the run, which is the same outcome as main. Spec 3 would serve stable but break every v5 target. So this PR makes misses rare (10 s first budget), cheap to repeat (cache, 2 s after a miss) and visible (span attributes and a warning), and keeps 6. Once [world-vercel] Attest the executor's spec version on run_started #4366 and vercel/workflow-server#1044 are deployed, a v5 executor raises an under-stamped run on run_started, so the miss floor can drop to 3 with a 2 s budget. That is tracked in Lower the cross-deployment probe-miss floor to spec 3 once executors raise under-stamped runs #4401.
  • Same-deployment starts are byte-for-byte unchanged, and opts.specVersion still overrides everything, so existing callers that pass a version keep working.

Alternatives considered

  1. Let the executor's run_started settle the version, treating run_created's as provisional (option 2 in the issue). As a general mechanism it is rejected because it changes every World and the workflow-server, and it is too late for choices the caller commits first (argument compression, CBOR vs JSON transport, attribute seeding). A narrow form of it is being done separately and complements this PR: [world-vercel] Attest the executor's spec version on run_started #4366 (the executor attests executorSpecVersion on run_started) and vercel/workflow-server#1044 (the backend raises a run to the executor's full version). Crossing a structural version (slot identity 6, the sealed log 7, or any unclassified future one) needs a fresh log; raises across capability-only versions (3, 4, 5, 8) apply any time before the run finishes. That is the only fix that reaches v4 callers already deployed, which no caller-side change can. The two follow different rules: this PR caps at the caller, and Enable lazyDiscovery in workbench #1044 raises the run to the executor's version after the creation round-trip (or at creation, on resilient start).
  2. Per-gate attestation in executionContext (for example hookForceClaimReaderVersion, reverted in 0feb2fd on hook-force-claim). Rejected so there is one versioning mechanism. Every gate would need its own attestation, while fixing the stamp fixes all gates at once.
  3. Use the probed version uncapped (what wf inspect's replay did before this PR). Rejected because a newer target could make an older caller stamp a version it cannot write or read back.
  4. Fall back to the caller's version on a probe miss (status quo). Rejected because that is the worst available answer when the target is older: it is exactly the over-stamp that hangs runs.
  5. Fail start() when the probe misses. Rejected because it would break starts that work today: same-major targets on a miss are served correctly by the spec-6 floor, and targets predating the health check never answer.
  6. Floor probed values at SPEC_VERSION_SUPPORTS_SLOT_IDENTITY as well (suggested in triage). Rejected because a target that reports 3 (a stable 4.x deployment) cannot read a 6-stamped run. A reported value is trusted even when it is below the runtime floor. The floor applies only when we know nothing.
  7. Cap at SPEC_VERSION_MAX_SUPPORTED instead of world.specVersion. Rejected, narrowly. The two are equal today. Capping at the World's minted version also respects the caller's own WORKFLOW_SEALED_LOG kill switch, which is the more conservative choice.
  8. Retry the probe on a miss (suggested in review). Rejected in favour of one probe with a longer budget. healthCheck() already returns on the first answer, so a longer single wait costs the same on a hit, and on a miss it does not discard a late answer to the first probe or send a second queue message.
  9. Read the target's version from deployment metadata recorded at build time, instead of a live probe. That would avoid waiting on a cold start, but it needs a new platform API and does not work for other Worlds. It is a possible follow-up.
  10. Sequence [world-vercel] Attest the executor's spec version on run_started #4366 / vercel/workflow-server#1044 first and keep 2 s with a floor of 3 (suggested in review). Not done here because main now mints 8, so the over-stamp against 7/7 targets is live today and an answered probe fixes it. The per-deployment cache keeps the 10 s budget off the steady state in the meantime, and Lower the cross-deployment probe-miss floor to spec 3 once executors raise under-stamped runs #4401 tracks the switch.

Not in this PR

  • Mixed-version tests for the reader gates themselves (sealed-log noop, hook force-claim). Those gates read run.specVersion, so these tests pin the stamp they read. The force-claim gate (feat(hooks): createHook({ experimental_force: true }) takes a held token over #4193) is now on main; its reader test lives with it. The new tests cover new starter → older executor, older starter → newer executor, pre-CBOR and pre-JSON targets, probe misses, explicit override, and target-named attribute errors.

Tests

  • packages/core/src/runtime/start-cross-deployment-spec-version.test.ts (new, 27 tests). Since the last round: an 8/8 caller and a 7/7 target stamp 7 (the case live on main), a malformed JSON reply stamps 3 with CBOR transport, an answer is reused for a second start to the same deployment, an answer after a miss is cached, and the second miss waits only the short budget. Besides the original cases, it covers a cold stable (spec 3) target answering after 5 s, which is stamped 3 (this test fails with the old 2 s budget), and an end-to-end probe timeout (stamped 6 in all three places, warns once across two starts, records probe-miss and the error). It also covers deploymentId: 'latest' resolving to another deployment and to the caller, retention errors naming the target, a probe below 5 disabling compression even when the core version says gzip, span attribution of an explicit version, and the resolver's source for every branch.
  • packages/core/src/runtime/runs.test.ts: three new recreateRunFromExisting cases (same-deployment pin, redirected replay left unset, explicit override).
  • On the branch merged with main (spec 8): packages/core src/ 2703 passed, the CLI inspect tests 120 passed, pnpm --filter @workflow/core typecheck, and biome check are clean. The only warnings are two pre-existing complexity warnings.

Triage and repro: https://taskmaster.playground-vercel.tools/s/ses_2ruvkbsezfsj

🤖 Generated with Claude Code

`start({ deploymentId })` targeting another deployment stamped the run
with the caller World's spec version, even though the target executes
it. Resolve it from the capability probe `start()` already awaits,
capped at the caller World's version; on a probe miss fall back to the
lowest version the runtime serves rather than the caller's. Redirected
`recreateRunFromExisting` replays no longer pin the source run's
version.

Closes #4251

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>

Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>
@changeset-bot

changeset-bot Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 29197a1

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
Name Type
@workflow/core Patch
@workflow/cli Patch
@workflow/builders Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
example-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workflow-docs Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC
workflow-web Ready Ready Preview, v0 Sep 25, 2026 6:18pm UTC

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 29197a1 · Fri, 25 Sep 2026 18:42:53 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 385 (-78%) 💚 2433 🔴 (+29%) 🔻 2499 🔴 (+30%) 🔻 2904 🔴 (+37%) 🔻 30
TTFS stream 294 (-76%) 💚 2591 🔴 (+46%) 🔻 2856 🔴 (+56%) 🔻 2947 🔴 (+55%) 🔻 30
TTFS hook + stream 582 (+7.8%) 3293 🔴 (+112%) 🔻 3392 🔴 (+102%) 🔻 3438 🔴 (-50%) 💚 30
Fan-out TTFS Promise.all(100 steps) 645 (+21%) 🔻 793 (+8.8%) 1056 (+42%) 🔻 2870 (+70%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 2428 (±0%) 5918 (+94%) 🔻 6175 (+91%) 🔻 9749 (+92%) 🔻 10
STSO 1020 steps (inline) 132 (+3.9%) 166 (+0.6%) 183 (-1.1%) 289 (-18%) 💚 1019
WO 1020 steps 164760 (±0%) 164760 (±0%) 164760 (±0%) 164760 (±0%) 1
CRTT first chunk (pooled) 83 (+12%) 121 (+8.0%) 148 (-18%) 💚 202 (-45%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 103 (±0%) 171 (-6%) 243 (-1%) 328 (-25%) 139 (-3%) 10
size sweep (100/s, 160B-12KB) 101 (+5%) 175 (-13%) 232 (-25%) 349 (-35%) 170 (-5%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 99 (+4%) 151 (-61%) 179 (-93%) 435 (-89%) 472 (+82%) 3
replay eve-gpt-5.6-sol-2000t (1x) 134 (+31%) 163 (-7%) 218 (±0%) 456 (-72%) 475 (-53%) 2
replay eve-gpt-5.6-sol-2000t (2x) 141 (+72%) 228 (+2%) 317 (+11%) 493 (-29%) 207 (-46%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 164957ms → this run 164332ms (Δ -625ms, 0%)

  100-150 ms  ██████████████┃████       main 480  this 369  -111
  150-200 ms  ███████████████████░░░░┃  main 474  this 595  +121
  200-250 ms  ┃█                        main  38  this  32    -6
  250-300 ms  ┃                         main  14  this  15    +1
  300-350 ms  ┃                         main   1  this   5    +4
  350-400 ms  ┃                         main   5  this   1    -4
  400-450 ms  ┃                         main   2  this   0    -2
  450-500 ms  ┃                         main   1  this   1    +0
  500-550 ms  ┃                         main   1  this   0    -1
  550-600 ms  ┃                         main   0  this   1    +1
  900-950 ms  ┃                         main   2  this   0    -2
1050-1100 ms  ┃                         main   1  this   0    -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▂█▂····   141.6 (-6%)   133 (-9%)   243 (-1%)  328 (-25%)  3000
sweep    ······▂█▂····  142.4 (-13%)  129 (-12%)  232 (-25%)  349 (-35%)  3000
gw 1x    ······▃█▁▁···  134.3 (-59%)  129 (-18%)  179 (-93%)  435 (-89%)  5295
eve 1x   ······▃█▂▁···  144.2 (-18%)  126 (-14%)   218 (±0%)  456 (-72%)  5186
eve 2x   ······▁█▄▁···   179.1 (-5%)   168 (-3%)  317 (+11%)  493 (-29%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  █▅▂▃▁▃▃▁▂▁  129–176ms
sweep    █▃▃▃▄▄▄▂▁▅  130–163ms
gw 1x    █▂▂▂▁▃▅▁▁▂  124–166ms
eve 1x   █▁▃▆▃▇▄▆▄▅  123–163ms
eve 2x   ▂▃▂▁▁▃▅▇█▄  152–222ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▅█▇█▅▄▁  139–145ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  █▅▁▅▁▅▃▄▅▄  41–53ms
sweep    ▁▄▆▅▅▆▄▂▃█  55–72ms
gw 1x    █▃▄▁▂▄▄▄▂▅  34–57ms
eve 1x   █▁▁▄▂▃▁▃▂▄  24–36ms
eve 2x   ▇▇▁▂▂▁▃▂▄█  24–34ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

❌ Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (62 failed)

python-node (62 failed):

  • promiseAllWorkflow | wrun_41M3CWP4N00GRB8F81TGZ0TARA | 🔍 observability
  • promiseRaceWorkflow | wrun_41M3CWP4N00GRB8F81TGZ0TARB | 🔍 observability
  • promiseAnyWorkflow | wrun_41M3CWP4N00GRB8F81TGZ0TARC | 🔍 observability
  • retainedInterleavingWorkflow | wrun_41M3CWPAY60GGNBFYK3PF54Y10 | 🔍 observability
  • hookWorkflow | wrun_41M3CWPH0V0GHYR9V7NKNN1YPB | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41M3CWPHRP0GYHYP7NDB8C0MQR | 🔍 observability
  • sleepingWorkflow | wrun_41M3CWPSN40GQP3RXB1CM8HWKM | 🔍 observability
  • parallelSleepWorkflow | wrun_41M3CWPTD80GRMG7HNTQ5HBP10 | 🔍 observability
  • sleepWinsRaceWorkflow | wrun_41M3CWPVQZ0GP2TXZA5K91W438 | 🔍 observability
  • stepWinsRaceWorkflow | wrun_41M3CWPZ4M0GT67BT349EQ3045 | 🔍 observability
  • nullByteWorkflow | wrun_41M3CWQ0K30GR77XXT2EHE0EYQ | 🔍 observability
  • outputStreamWorkflow no startIndex (reads all chunks)
  • outputStreamWorkflow positive startIndex (skips first chunk)
  • outputStreamWorkflow negative startIndex (reads from end)
  • outputStreamWorkflow - getTailIndex and getChunks getTailIndex returns correct index after stream completes
  • outputStreamWorkflow - getTailIndex and getChunks getChunks returns same content as reading the stream
  • outputStreamInsideStepWorkflow - getWritable() called inside step functions | wrun_41M3CWQMQH0GGMNYN2SKP80V0Y | 🔍 observability
  • utf8StreamWorkflow | wrun_41M3CWQSYY0GKEG36KPPXWKASJ | 🔍 observability
  • writableForwardedFromWorkflowWorkflow | wrun_41M3CWQTGX0GXNEAV0JCS0T6MA | 🔍 observability
  • writableForwardedFromStepWorkflow | wrun_41M3CWQTHT0GZ6RK46CKG4EWRJ | 🔍 observability
  • promiseRaceStressTestWorkflow | wrun_41M3CWQXB60GH2XDKT3H1K0PP3 | 🔍 observability
  • error handling retry behavior regular Error retries until success
  • error handling retry behavior FatalError fails immediately without retries
  • error handling retry behavior RetryableError respects custom retryAfter delay
  • error handling retry behavior maxRetries=0 disables retries
  • error handling catchability FatalError can be caught and detected with FatalError.is()
  • error handling catchability step throw round-trips FatalError with cause chain to workflow catch
  • error handling catchability workflow throw round-trips FatalError + cause through run_failed event
  • error handling not registered StepNotRegisteredError fails the step but workflow can catch it
  • error handling not registered StepNotRegisteredError fails the run when not caught in workflow
  • hookCleanupTestWorkflow - hook token reuse after workflow completion | wrun_41M3CWSNGQ0GJDAW79TK2TXVEQ | 🔍 observability
  • concurrent hook token conflict - two workflows cannot use the same hook token simultaneously | wrun_41M3CWSFAR0GX1PAWYG82S54TJ | 🔍 observability
  • hookGetConflictWorkflow - awaiting hook.getConflict() registers hook without payload | wrun_41M3CWSKBS0GHY1P8ZVS7Y8NCK | 🔍 observability
  • 'hookGetConflictWithPriorStepWorkflow' - hook.getConflict() does not block step execution | wrun_41M3CWSQMY0GHN57V11MA0MYG8 | 🔍 observability
  • 'hookGetConflictWithParallelStepWorkfl…' - hook.getConflict() does not block step execution | wrun_41M3CWSSR80GYMHXWPMDY0QZVM | 🔍 observability
  • hookGetConflictThenStepParallelWorkflow - hook.getConflict() continuation step runs alongside other steps | wrun_41M3CWSVVX0GZKE85JEE3W90G2 | 🔍 observability
  • hookGetConflictWorkflow - hook.getConflict() resolves with the conflicting run when token is already registered | wrun_41M3CWSYYZ0GVB07RA64MQ1MXV | 🔍 observability
  • hookClaimOnlyMutexWorkflow - hook works as a pure run mutex without payload data | wrun_41M3CWTJ1X0GG7YRMQR4XS68YA | 🔍 observability
  • hookAdoptOwnerResultWorkflow - duplicate adopts the owner result via conflict.returnValue | wrun_41M3CWT6BP0GTHB5K7P5A9MQMQ | 🔍 observability
  • hookSignalOwnerWorkflow - duplicate forwards its payload to the owner via resumeHook | wrun_41M3CWTCQQ0GPDSMBWKMMG99ZT | 🔍 observability
  • resume-or-start route pattern - resumeHook retried after start() reaches the new run | wrun_41M3CWV2ZQ0GYWRJFW19RMD1QD | 🔍 observability
  • hookDisposeTestWorkflow - hook token reuse after explicit disposal while workflow still running | wrun_41M3CWV7FC0GZ92RD8KJHCK9CH | 🔍 observability
  • hookTokenReuseLoopWorkflow - same run recreates a hook with the same token after dispose() | wrun_41M3CWV7RA0GTKK4AKNYDHTSQQ | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_41M3CWVS1G0GSYF6D207K15234 | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M3CWWRFE0GW4Z5PMF3BW2WM3 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M3CWWRNY0GT3RESYNYF8XJ1C | 🔍 observability
  • hookWithSleepWorkflow - hook payloads delivered correctly with concurrent sleep | wrun_41M3CWWXDH0GP31ZPHZ3MTEE5E | 🔍 observability
  • hookWithSleepFinalStepWorkflow - step only on final payload | wrun_41M3CWX1MM0GTKF8F1K1VXZYFR | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M3CWX2SC0GY5DN3Y9S8Y32H0 | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_41M3CWX31F0GQXB6J1TM37MEMG | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
  • AbortController abortParallelWorkflow: abort cancels all parallel steps
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortViaHookWorkflow: external hook triggers abort on in-flight step
  • metadataFromHelperWorkflow - getWorkflowMetadata/getStepMetadata work from module-level helper (#1577) | wrun_41M3CWXHA50GQT79RVYB2MYYKR | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M3CWXJ5X0GYW484WRJDP9ATV | 🔍 observability
  • setAttributes start: initial attributes are seeded on run creation
  • setAttributes start: reserved-prefix initial attributes are seeded with allowReservedAttributes
  • setAttributes setAttributesWorkflow: workflow-body calls append native attr_set events and merge correctly
  • setAttributes setAttributesInsideStepWorkflow: step-body calls append attributed native events
  • setAttributes Promise.all of disjoint-key writes: every key lands
  • setAttributes workflow throws after awaited setAttributes: attribute still persists on the failed run

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • a chain of takeovers: each victim ends with HookForceClaimedError, deliveries follow the current owner (nextjs-webpack · local-dev / local / node / stable)
  • health check (CLI) - workflow health command reports healthy endpoints (nextjs-webpack · local-dev / local / quickjs / canary)
  • resume-or-start route pattern - resumeHook retried after start() reaches the new run (nextjs-webpack · local-dev / local / node / stable)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

125 infra events
  • cold-start-warmup · suite warmup (tanstack-start) · at 18:20:52Z · abandoned wrun_01M3CWRD0XX2H3QBV2E6YW1FR2
  • run-pickup-stall · hookMinRetentionWorkflow - terminal Hook cannot resume and its token stays unavailable (nuxt) · at 18:22:47Z · abandoned wrun_01M3CWW0KH5MWZPE7AG601MWS7
  • run-pickup-stall · WorkflowNotRegisteredError fails the run when workflow does not exist (nextjs-webpack) · at 18:26:10Z · abandoned wrun_01M3CX273R42XF3M1D1H55M8EP
  • cold-start-warmup · suite warmup (python) · at 18:35:59Z · abandoned wrun_41M3CXGZF10GR0YQT5371J6H7P · (+7 more)
  • run-pickup-stall · hookWorkflow (python) · at 18:36:16Z · abandoned wrun_41M3CXMMY40GVVXD422WRTMJHF
  • run-pickup-stall · promiseRaceWorkflow (python) · at 18:36:16Z · abandoned wrun_41M3CXMMY10GRSF4NWZ4TM9KQ6
  • run-pickup-stall · promiseAllWorkflow (python) · at 18:36:16Z · abandoned wrun_41M3CXMMY10GRSF4NWZ4TM9KQ5
  • run-pickup-stall · promiseAnyWorkflow (python) · at 18:36:16Z · abandoned wrun_41M3CXMMY10GRSF4NWZ4TM9KQ7
  • run-pickup-stall · retainedInterleavingWorkflow (python) · at 18:36:17Z · abandoned wrun_41M3CXMMY10GRSF4NWZ4TM9KQ8
  • run-pickup-stall · hookWorkflow (python) · at 18:37:04Z · abandoned wrun_41M3CXP48C0GNG46T7P4TT9K1G
  • run-pickup-stall · retainedInterleavingWorkflow (python) · at 18:37:04Z · abandoned wrun_41M3CXP55C0GXDM5Z94VNSPJVH
  • run-pickup-stall · promiseAnyWorkflow (python) · at 18:37:16Z · abandoned wrun_41M3CXPGET0GXCV524YBKAE1FP
  • run-pickup-stall · promiseAllWorkflow (python) · at 18:37:16Z · abandoned wrun_41M3CXPGEH0GW2XGC432XCT8ZF
  • run-pickup-stall · promiseRaceWorkflow (python) · at 18:37:17Z · abandoned wrun_41M3CXPGF20GNDMYVYZDE34TB8
  • run-pickup-stall · hookWorkflow is not resumable via public webhook endpoint (python) · at 18:37:52Z · abandoned wrun_41M3CXQJWD0GMDNBB4GXKXV59Z
  • run-pickup-stall · sleepingWorkflow (python) · at 18:37:53Z · abandoned wrun_41M3CXQMBP0GMMT1FZ18ESNYHG
  • run-pickup-stall · parallelSleepWorkflow (python) · at 18:38:19Z · abandoned wrun_41M3CXRCKK0GQRJCW00VRP905S
  • run-pickup-stall · sleepWinsRaceWorkflow (python) · at 18:38:19Z · abandoned wrun_41M3CXRD3C0GMWGTP9J5HW5DM1
  • run-pickup-stall · stepWinsRaceWorkflow (python) · at 18:38:44Z · abandoned wrun_41M3CXRD9D0GG5JB7EHP9WGK0M
  • run-pickup-stall · hookWorkflow is not resumable via public webhook endpoint (python) · at 18:39:54Z · abandoned wrun_41M3CXT3FX0GXC405PAPXE0KBY
  • run-pickup-stall · sleepingWorkflow (python) · at 18:39:54Z · abandoned wrun_41M3CXT3PD0GWAA6BK1B0DEVMV
  • run-pickup-stall · parallelSleepWorkflow (python) · at 18:39:54Z · abandoned wrun_41M3CXV4G40GS10AKRZ6E4S525
  • run-pickup-stall · sleepWinsRaceWorkflow (python) · at 18:39:55Z · abandoned wrun_41M3CXV4ZY0GRPFXWJ7SDKCA2T
  • run-pickup-stall · stepWinsRaceWorkflow (python) · at 18:39:55Z · abandoned wrun_41M3CXV55X0GX49RKFAZNN56SG
  • run-pickup-stall · nullByteWorkflow (python) · at 18:40:15Z · abandoned wrun_41M3CXVYXC0GHPWVDW21B2XD6D
  • run-pickup-stall · no startIndex (reads all chunks) (python) · at 18:40:15Z · abandoned wrun_41M3CXVZ250GY1NHM3R781T8MG
  • run-pickup-stall · positive startIndex (skips first chunk) (python) · at 18:40:48Z · abandoned wrun_41M3CXWZ350GJBF9D6S46R1N09
  • run-pickup-stall · negative startIndex (reads from end) (python) · at 18:40:48Z · abandoned wrun_41M3CXWZJY0GYPWH2BEYB1HVXK
  • run-pickup-stall · getTailIndex returns correct index after stream completes (python) · at 18:40:51Z · abandoned wrun_41M3CXX28G0GP6SDKXA8JTWJ0G
  • run-pickup-stall · getTailIndex returns -1 before any chunks are written (python) · at 18:41:15Z · abandoned wrun_41M3CXXSGD0GMAV59HRSMKHRWX
  • run-pickup-stall · getChunks returns same content as reading the stream (python) · at 18:41:15Z · abandoned wrun_41M3CXXSN40GZX56AKAA8809TH
  • run-pickup-stall · outputStreamInsideStepWorkflow - getWritable() called inside step functions (python) · at 18:41:31Z · abandoned wrun_41M3CXY9R60GY78EY2EJW3EA7X
  • run-pickup-stall · utf8StreamWorkflow (python) · at 18:41:48Z · abandoned wrun_41M3CXYSP40GQ0KFSBBRTHZZE1
  • run-pickup-stall · nullByteWorkflow (python) · at 18:41:51Z · abandoned wrun_41M3CXYWVG0GM9TVKEGDEP7Z47
  • run-pickup-stall · no startIndex (reads all chunks) (python) · at 18:41:51Z · abandoned wrun_41M3CXYXJP0GYXNEV93WPW5GQJ
  • run-pickup-stall · positive startIndex (skips first chunk) (python) · at 18:42:15Z · abandoned wrun_41M3CXZM850GQ93FJKZFK4XWF0
  • run-pickup-stall · negative startIndex (reads from end) (python) · at 18:42:32Z · abandoned wrun_41M3CY054H0GR5CVW8AVBPHY16
  • run-pickup-stall · getTailIndex returns correct index after stream completes (python) · at 18:42:52Z · abandoned wrun_41M3CY0R7D0GVW3F8DJTK3PGWF
  • run-pickup-stall · getChunks returns same content as reading the stream (python) · at 18:42:54Z · abandoned wrun_41M3CY0TQ20GVWY713BMW4YR5X
  • run-pickup-stall · outputStreamInsideStepWorkflow - getWritable() called inside step functions (python) · at 18:43:15Z · abandoned wrun_41M3CY1EV50GN6HNYMW197T7GB
  • run-pickup-stall · writableForwardedFromWorkflowWorkflow (python) · at 18:43:32Z · abandoned wrun_41M3CY1ZQH0GTC3TV5K3MBT8HZ
  • run-pickup-stall · writableForwardedFromStepWorkflow (python) · at 18:43:50Z · abandoned wrun_41M3CY2HEG0GX3KTAGTDG2MRNA
  • run-pickup-stall · utf8StreamWorkflow (python) · at 18:43:51Z · abandoned wrun_41M3CY2JTD0GTQD3DXYG3K4BRV
  • run-pickup-stall · promiseRaceStressTestWorkflow (python) · at 18:43:56Z · abandoned wrun_41M3CY2Q4G0GWDY97M081JJN0A
  • run-pickup-stall · regular Error retries until success (python) · at 18:44:16Z · abandoned wrun_41M3CY3A7E0GKNVB7RBZVS6F3H
  • run-pickup-stall · FatalError fails immediately without retries (python) · at 18:44:57Z · abandoned wrun_41M3CY4HQH0GHC28KD5597KGP2
  • run-pickup-stall · RetryableError respects custom retryAfter delay (python) · at 18:45:16Z · abandoned wrun_41M3CY54TD0GJVWJ5W7513Y4KG
  • run-pickup-stall · maxRetries=0 disables retries (python) · at 18:45:32Z · abandoned wrun_41M3CY5MXH0GWTH6VCBVFYWYN4
  • run-pickup-stall · FatalError can be caught and detected with FatalError.is() (python) · at 18:45:50Z · abandoned wrun_41M3CY66MH0GV38ZHYPVHY5TZE
  • run-pickup-stall · step throw round-trips FatalError with cause chain to workflow catch (python) · at 18:45:52Z · abandoned wrun_41M3CY680E0GW2CMAYNT5GQBGC
  • run-pickup-stall · workflow throw round-trips FatalError + cause through run_failed event (python) · at 18:45:56Z · abandoned wrun_41M3CY6CAJ0GZ8RR0EWJXAQ8ZN
  • run-pickup-stall · StepNotRegisteredError fails the step but workflow can catch it (python) · at 18:46:16Z · abandoned wrun_41M3CY6ZDF0GQ5C78H9M1RBGG8
  • run-pickup-stall · StepNotRegisteredError fails the run when not caught in workflow (python) · at 18:46:39Z · abandoned wrun_41M3CY7FGJ0GVKG1VP4SXF80A2
  • run-pickup-stall · promiseRaceStressTestWorkflow (python) · at 18:47:45Z · abandoned wrun_41M3CY96R90GM7VEAXC626CHVW
  • run-pickup-stall · writableForwardedFromWorkflowWorkflow (python) · at 18:48:10Z · abandoned wrun_41M3CY9CCR0GSX99KJWSJZ7Q2N
  • run-pickup-stall · writableForwardedFromStepWorkflow (python) · at 18:48:12Z · abandoned wrun_41M3CY9CYP0GY5GGVD6WX2BTKQ
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (python) · at 18:48:12Z · abandoned wrun_41M3CY9TX00GR56B419JBB4WBB
  • run-pickup-stall · regular Error retries until success (python) · at 18:48:28Z · abandoned wrun_41M3CY9A3J0GNWQT00Q8Q7W0WC
  • run-pickup-stall · FatalError fails immediately without retries (python) · at 18:49:12Z · abandoned wrun_41M3CYB4PK0GVQM5F1X1ZDJKW3
  • run-pickup-stall · RetryableError respects custom retryAfter delay (python) · at 18:49:22Z · abandoned wrun_41M3CYBNG00GSFRDF20T1DDEYN
  • run-pickup-stall · maxRetries=0 disables retries (python) · at 18:49:28Z · abandoned wrun_41M3CYBYMT0GW0RVG8BF6MCSCY
  • run-pickup-stall · FatalError can be caught and detected with FatalError.is() (python) · at 18:49:40Z · abandoned wrun_41M3CYD1JS0GN0BQGDNETK3CPY
  • run-pickup-stall · step throw round-trips FatalError with cause chain to workflow catch (python) · at 18:49:41Z · abandoned wrun_41M3CYD24P0GTF5B6DTRJAF3SJ
  • run-pickup-stall · workflow throw round-trips FatalError + cause through run_failed event (python) · at 18:49:45Z · abandoned wrun_41M3CYDAVA0GJ09JGMMPBHKSBD
  • run-pickup-stall · StepNotRegisteredError fails the step but workflow can catch it (python) · at 18:49:52Z · abandoned wrun_41M3CYDJFE0GJSA465ZNMPZB55
  • run-pickup-stall · StepNotRegisteredError fails the run when not caught in workflow (python) · at 18:50:01Z · abandoned wrun_41M3CYDVNZ0GVD5WK9ZV4Y26CH
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (python) · at 18:50:35Z · abandoned wrun_41M3CYEWYQ0GKCSTKS8692358R
  • run-pickup-stall · hookGetConflictWorkflow - awaiting hook.getConflict() registers hook without payload (python) · at 18:50:45Z · abandoned wrun_41M3CYF6BD0GSCKQDBD8YZ8HHJ
  • run-pickup-stall · 'hookGetConflictWithPriorStepWorkflow' - hook.getConflict() does not block step execution (python) · at 18:50:52Z · abandoned wrun_41M3CYFD2E0GVY3R05AHKGQZ5J
  • run-pickup-stall · 'hookGetConflictWithParallelStepWorkfl…' - hook.getConflict() does not block step execution (python) · at 18:51:05Z · abandoned wrun_41M3CYFT840GY1R0BF3MQT52KD
  • run-pickup-stall · hookGetConflictThenStepParallelWorkflow - hook.getConflict() continuation step runs alongside other steps (python) · at 18:51:22Z · abandoned wrun_41M3CYGAXY0GN32YXQ14WA7XM0
  • run-pickup-stall · hookGetConflictWorkflow - hook.getConflict() resolves with the conflicting run when token is already registered (python) · at 18:51:37Z · abandoned wrun_41M3CYGSJR0GHCGPS35HN8B9E3
  • run-pickup-stall · hookGetConflictWorkflow - awaiting hook.getConflict() registers hook without payload (python) · at 18:51:46Z · abandoned wrun_41M3CYH1S00GM2K38PMS32176Q
  • run-pickup-stall · 'hookGetConflictWithPriorStepWorkflow' - hook.getConflict() does not block step execution (python) · at 18:51:53Z · abandoned wrun_41M3CYH8DY0GZSEQ3N12AQEA1Y
  • run-pickup-stall · 'hookGetConflictWithParallelStepWorkfl…' - hook.getConflict() does not block step execution (python) · at 18:52:07Z · abandoned wrun_41M3CYHNKA0GMH4RGGENM0T18B
  • run-pickup-stall · hookGetConflictWorkflow - hook.getConflict() resolves with the conflicting run when token is already registered (python) · at 18:52:26Z · abandoned wrun_41M3CYJ8340GK5PME5XP7JY0PP
  • run-pickup-stall · hookClaimOnlyMutexWorkflow - hook works as a pure run mutex without payload data (python) · at 18:52:46Z · abandoned wrun_41M3CYJX5N0GHTPMR5KGDMHRYG
  • run-pickup-stall · hookGetConflictThenStepParallelWorkflow - hook.getConflict() continuation step runs alongside other steps (python) · at 18:52:53Z · abandoned wrun_41M3CYK3PE0GKX30B79E47232H
  • run-pickup-stall · hookAdoptOwnerResultWorkflow - duplicate adopts the owner result via conflict.returnValue (python) · at 18:52:53Z · abandoned wrun_41M3CYK40A0GW4YYD48ZK5DF37
  • run-pickup-stall · hookSignalOwnerWorkflow - duplicate forwards its payload to the owner via resumeHook (python) · at 18:53:07Z · abandoned wrun_41M3CYKGZ00GKG3YHHYP4X5RPD
  • run-pickup-stall · resume-or-start route pattern - resumeHook retried after start() reaches the new run (python) · at 18:53:14Z · abandoned wrun_41M3CYKQKR0GYQAATT3JVNDM4Z
  • run-pickup-stall · hookClaimOnlyMutexWorkflow - hook works as a pure run mutex without payload data (python) · at 18:53:34Z · abandoned wrun_41M3CYMBDM0GX8MZX5SJ4YHR4S
  • run-pickup-stall · hookAdoptOwnerResultWorkflow - duplicate adopts the owner result via conflict.returnValue (python) · at 18:53:44Z · abandoned wrun_41M3CYMMYG0GT3XRW26QABRZ2B
  • run-pickup-stall · hookSignalOwnerWorkflow - duplicate forwards its payload to the owner via resumeHook (python) · at 18:53:54Z · abandoned wrun_41M3CYMYX60GPX2PX0RC0G6M9T
  • run-pickup-stall · resume-or-start route pattern - resumeHook retried after start() reaches the new run (python) · at 18:54:01Z · abandoned wrun_41M3CYN6320GHR0ERV6JEA83BN
  • run-pickup-stall · hookDisposeTestWorkflow - hook token reuse after explicit disposal while workflow still running (python) · at 18:54:21Z · abandoned wrun_41M3CYNSKM0GMREJXPFMSMQVB1
  • run-pickup-stall · hookTokenReuseLoopWorkflow - same run recreates a hook with the same token after dispose() (python) · at 18:54:24Z · abandoned wrun_41M3CYNW9X0GH97E6V4XHYGZX4
  • run-pickup-stall · spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step (python) · at 18:54:31Z · abandoned wrun_41M3CYP38K0GSR0BMF6KCMZTAV
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 18:54:45Z · abandoned wrun_41M3CYPHAB0GRZWBSB3G39TF2B
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 18:54:48Z · abandoned wrun_41M3CYPM9R0GKWSJ0QXHVCEXKV
  • run-pickup-stall · hookDisposeTestWorkflow - hook token reuse after explicit disposal while workflow still running (python) · at 18:55:08Z · abandoned wrun_41M3CYQ7CM0GZ7WNSC1TXC9X3J
  • run-pickup-stall · hookTokenReuseLoopWorkflow - same run recreates a hook with the same token after dispose() (python) · at 18:55:11Z · abandoned wrun_41M3CYQAHV0GP9QKZQ35YC95KY
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 18:55:18Z · abandoned wrun_41M3CYQGZS0GQ268G6Z95VDBCM
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 18:55:21Z · abandoned wrun_41M3CYQKN70GXX98NCTH1KXA4B
  • run-pickup-stall · hookWithSleepWorkflow - hook payloads delivered correctly with concurrent sleep (python) · at 18:55:50Z · abandoned wrun_41M3CYRGES0GT768J4NSH699JM
  • run-pickup-stall · hookWithSleepFinalStepWorkflow - step only on final payload (python) · at 18:55:53Z · abandoned wrun_41M3CYRJZC0GRQX70RJ820A1D3
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 18:55:55Z · abandoned wrun_41M3CYRNKD0GWCWFX5GB326VG2
  • run-pickup-stall · sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) (python) · at 18:55:59Z · abandoned wrun_41M3CYRRPF0GX6MTFVDB8QA4EV
  • run-pickup-stall · spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step (python) · at 18:56:32Z · abandoned wrun_41M3CYSSBJ0GS89A60VYCRQDV2
  • run-pickup-stall · hookWithSleepWorkflow - hook payloads delivered correctly with concurrent sleep (python) · at 18:56:37Z · abandoned wrun_41M3CYSYKB0GTRHSQRSGB1AS3R
  • run-pickup-stall · hookWithSleepFinalStepWorkflow - step only on final payload (python) · at 18:56:40Z · abandoned wrun_41M3CYT0XS0GKF55M9PCMV1ZQQ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 18:56:56Z · abandoned wrun_41M3CYTGZB0GJJHM1AYMPE37WN
  • run-pickup-stall · sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) (python) · at 18:57:01Z · abandoned wrun_41M3CYTMBQ0GZSS7J6BF70HN8G
  • run-pickup-stall · abortTimeoutWorkflow: timeout cancels long-running step (python) · at 18:57:40Z · abandoned wrun_41M3CYVNPF0GTH5FQYNJ1WFR93
  • run-pickup-stall · abortParallelWorkflow: abort cancels all parallel steps (python) · at 18:57:41Z · abandoned wrun_41M3CYVPNB0GMM4FKX74KMHHZC
  • run-pickup-stall · abortReasonWorkflow: abort reason preserved across boundaries (python) · at 18:58:21Z · abandoned wrun_41M3CYWBJB0GGASK4K9A8Q0YVP
  • run-pickup-stall · metadataFromHelperWorkflow - getWorkflowMetadata/getStepMetadata work from module-level helper (Vercel World Regression: getWorkflowMetadata() fully breaking runs #1577) (python) · at 18:58:22Z · abandoned wrun_41M3CYWEYR0GHX0PZN9WJ41FFK
  • run-pickup-stall · abortViaHookWorkflow: external hook triggers abort on in-flight step (python) · at 18:59:23Z · abandoned wrun_41M3CYXG9G0GKJPDWT2DF4NGDK
  • run-pickup-stall · abortTimeoutWorkflow: timeout cancels long-running step (python) · at 18:59:52Z · abandoned wrun_41M3CYZ6V90GYZG7HX9FWWBVE1
  • run-pickup-stall · abortParallelWorkflow: abort cancels all parallel steps (python) · at 18:59:53Z · abandoned wrun_41M3CYZAWF0GKJ8QTNSRKQD7QQ
  • run-pickup-stall · start: initial attributes are seeded on run creation (python) · at 18:59:57Z · abandoned wrun_41M3CYZRRC0GN3AVBBZXXSQDW0
  • run-pickup-stall · start: reserved-prefix initial attributes are seeded with allowReservedAttributes (python) · at 18:59:57Z · abandoned wrun_41M3CYZS7P0GZS56SQQCN4815Q
  • run-pickup-stall · setAttributesWorkflow: workflow-body calls append native attr_set events and merge correctly (python) · at 19:00:08Z · abandoned wrun_41M3CZ094X0GND3TF5RAYPF3AV
  • run-pickup-stall · setAttributesInsideStepWorkflow: step-body calls append attributed native events (python) · at 19:00:18Z · abandoned wrun_41M3CZ0P1W0GHM0Z1R6FK5ZEB5
  • run-pickup-stall · abortReasonWorkflow: abort reason preserved across boundaries (python) · at 19:00:20Z · abandoned wrun_41M3CZ0QK70GGVACQDS4HV9CD6
  • run-pickup-stall · metadataFromHelperWorkflow - getWorkflowMetadata/getStepMetadata work from module-level helper (Vercel World Regression: getWorkflowMetadata() fully breaking runs #1577) (python) · at 19:00:30Z · abandoned wrun_41M3CZ11E90GJAYHWSZ9DMVFTB
  • run-pickup-stall · Promise.all of disjoint-key writes: every key lands (python) · at 19:00:35Z · abandoned wrun_41M3CZ16R30GTGWXXQ3VQ0BN5G
  • run-pickup-stall · abortViaHookWorkflow: external hook triggers abort on in-flight step (python) · at 19:00:36Z · abandoned wrun_41M3CZ17690GTTTGMDCDN7ZYQ4
  • run-pickup-stall · start: initial attributes are seeded on run creation (python) · at 19:01:02Z · abandoned wrun_41M3CZ213R0GX19XVK9VXDZG56
  • run-pickup-stall · start: reserved-prefix initial attributes are seeded with allowReservedAttributes (python) · at 19:01:05Z · abandoned wrun_41M3CZ241J0GJJMH1FBYWP3C57
  • run-pickup-stall · setAttributesWorkflow: workflow-body calls append native attr_set events and merge correctly (python) · at 19:01:21Z · abandoned wrun_41M3CZ2KAS0GQE69GWBT6WKGMZ
  • run-pickup-stall · setAttributesInsideStepWorkflow: step-body calls append attributed native events (python) · at 19:01:24Z · abandoned wrun_41M3CZ2PDD0GM9Q8FSKMN5B05M
  • run-pickup-stall · Promise.all of disjoint-key writes: every key lands (python) · at 19:01:31Z · abandoned wrun_41M3CZ2WVE0GH3V7FWW1P7EVKP
  • run-pickup-stall · workflow throws after awaited setAttributes: attribute still persists on the failed run (python) · at 19:01:34Z · abandoned wrun_41M3CZ303X0GK89FZ4E9D669RV
  • run-pickup-stall · workflow throws after awaited setAttributes: attribute still persists on the failed run (python) · at 19:02:05Z · abandoned wrun_41M3CZ3Y6X0GYJW19SF0QE0438

E2E Test Summary

Summary
Passed Failed Skipped Total
❌ ▲ Vercel Production 3816 62 739 4617
✅ 💻 Local Development 4068 0 549 4617
✅ 📦 Local Production 4238 0 550 4788
✅ 🐘 Local Postgres 4238 0 550 4788
✅ 🪟 Windows 340 0 2 342
✅ vercel-http-transport 873 0 153 1026
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 591 0 93 684
Total 18191 62 2636 20889
Details by Category

❌ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 141 0 30
✅ astro-quickjs 141 0 30
✅ example-node 141 0 30
✅ example-quickjs 141 0 30
✅ express-node 141 0 30
✅ express-quickjs 141 0 30
✅ fastify-node 141 0 30
✅ fastify-quickjs 141 0 30
✅ hono-node 141 0 30
✅ hono-quickjs 141 0 30
✅ nest-node 141 0 30
✅ nest-quickjs 141 0 30
✅ nextjs-turbopack-node 168 0 3
✅ nextjs-turbopack-quickjs 168 0 3
✅ nextjs-webpack-node 168 0 3
✅ nextjs-webpack-quickjs 168 0 3
✅ nitro-node 141 0 30
✅ nitro-quickjs 141 0 30
✅ nuxt-node 141 0 30
✅ nuxt-quickjs 141 0 30
❌ python-node 4 62 105
✅ sveltekit-node 160 0 11
✅ sveltekit-quickjs 160 0 11
✅ tanstack-start-node 141 0 30
✅ tanstack-start-quickjs 141 0 30
✅ vite-node 141 0 30
✅ vite-quickjs 141 0 30

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 142 0 29
✅ astro-stable-quickjs 142 0 29
✅ express-stable-node 142 0 29
✅ express-stable-quickjs 142 0 29
✅ fastify-stable-node 142 0 29
✅ fastify-stable-quickjs 142 0 29
✅ hono-stable-node 142 0 29
✅ hono-stable-quickjs 142 0 29
✅ nest-stable-node 142 0 29
✅ nest-stable-quickjs 142 0 29
✅ nextjs-turbopack-canary-node 170 0 1
✅ nextjs-turbopack-canary-quickjs 170 0 1
✅ nextjs-turbopack-stable-node 170 0 1
✅ nextjs-turbopack-stable-quickjs 170 0 1
✅ nextjs-webpack-canary-quickjs 170 0 1
✅ nextjs-webpack-stable-node 170 0 1
✅ nextjs-webpack-stable-quickjs 170 0 1
✅ nitro-stable-node 142 0 29
✅ nitro-stable-quickjs 142 0 29
✅ nuxt-stable-node 142 0 29
✅ nuxt-stable-quickjs 142 0 29
✅ sveltekit-stable-node 161 0 10
✅ sveltekit-stable-quickjs 161 0 10
✅ tanstack-start-node 142 0 29
✅ tanstack-start-quickjs 142 0 29
✅ vite-stable-node 142 0 29
✅ vite-stable-quickjs 142 0 29

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 142 0 29
✅ astro-stable-quickjs 142 0 29
✅ express-stable-node 142 0 29
✅ express-stable-quickjs 142 0 29
✅ fastify-stable-node 142 0 29
✅ fastify-stable-quickjs 142 0 29
✅ hono-stable-node 142 0 29
✅ hono-stable-quickjs 142 0 29
✅ nest-stable-node 142 0 29
✅ nest-stable-quickjs 142 0 29
✅ nextjs-turbopack-canary-node 170 0 1
✅ nextjs-turbopack-canary-quickjs 170 0 1
✅ nextjs-turbopack-stable-node 170 0 1
✅ nextjs-turbopack-stable-quickjs 170 0 1
✅ nextjs-webpack-canary-node 170 0 1
✅ nextjs-webpack-canary-quickjs 170 0 1
✅ nextjs-webpack-stable-node 170 0 1
✅ nextjs-webpack-stable-quickjs 170 0 1
✅ nitro-stable-node 142 0 29
✅ nitro-stable-quickjs 142 0 29
✅ nuxt-stable-node 142 0 29
✅ nuxt-stable-quickjs 142 0 29
✅ sveltekit-stable-node 161 0 10
✅ sveltekit-stable-quickjs 161 0 10
✅ tanstack-start-node 142 0 29
✅ tanstack-start-quickjs 142 0 29
✅ vite-stable-node 142 0 29
✅ vite-stable-quickjs 142 0 29

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 142 0 29
✅ astro-stable-quickjs 142 0 29
✅ express-stable-node 142 0 29
✅ express-stable-quickjs 142 0 29
✅ fastify-stable-node 142 0 29
✅ fastify-stable-quickjs 142 0 29
✅ hono-stable-node 142 0 29
✅ hono-stable-quickjs 142 0 29
✅ nest-stable-node 142 0 29
✅ nest-stable-quickjs 142 0 29
✅ nextjs-turbopack-canary-node 170 0 1
✅ nextjs-turbopack-canary-quickjs 170 0 1
✅ nextjs-turbopack-stable-node 170 0 1
✅ nextjs-turbopack-stable-quickjs 170 0 1
✅ nextjs-webpack-canary-node 170 0 1
✅ nextjs-webpack-canary-quickjs 170 0 1
✅ nextjs-webpack-stable-node 170 0 1
✅ nextjs-webpack-stable-quickjs 170 0 1
✅ nitro-stable-node 142 0 29
✅ nitro-stable-quickjs 142 0 29
✅ nuxt-stable-node 142 0 29
✅ nuxt-stable-quickjs 142 0 29
✅ sveltekit-stable-node 161 0 10
✅ sveltekit-stable-quickjs 161 0 10
✅ tanstack-start-node 142 0 29
✅ tanstack-start-quickjs 142 0 29
✅ vite-stable-node 142 0 29
✅ vite-stable-quickjs 142 0 29

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 170 0 1
✅ nextjs-turbopack-quickjs 170 0 1

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 141 0 30
✅ express 141 0 30
✅ hono 141 0 30
✅ nextjs-turbopack 168 0 3
✅ nitro 141 0 30
✅ vite 141 0 30

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 141 0 30
✅ express 141 0 30
✅ nextjs-turbopack 168 0 3
✅ vite 141 0 30

📋 View full workflow run

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 42 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ step-vs-timer-early-settlement completed 8 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 265.3 KiB (+13.4 KiB) ⚠️ 96.0 KiB (+1.1 KiB) 1.93 MiB (+17.3 KiB)
nextjs-turbopack 272.8 KiB (+13.4 KiB) ⚠️ 426 B (±0) 922.5 KiB (+14.3 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

29197a1 · run

@kvnloo

kvnloo commented Sep 23, 2026

Copy link
Copy Markdown

one edge case i'm not following: on a probe timeout we fall back to spec 6, but the motivating case explicitly includes older cross-deployment targets like stable at spec 3.

so if that older target is healthy but cold enough to miss the 2s probe, don't we recreate the original over-stamp?

the caller can already mint lower-version runs — the pre-CBOR test demonstrates that — so i'm not sure the caller runtime's compatibility floor also needs to become the fallback for an unknown target.

would it be worth adding a regression for a new starter targeting a spec-3 deployment whose probe times out, and asserting that the fallback can't exceed what the unknown target may support?

…mped spec version

Address review on #4327:

- Raise the capability probe budget from 2s to 10s. healthCheck() returns
  on the first answer, so only a miss pays; a cold but healthy older-major
  target (e.g. stable, spec 3) is now read from the probe instead of being
  stamped with the spec-6 guess it would reject.
- Be honest about the miss floor: spec 6 is the lowest a v5 runtime
  executes, and is not safe for older-major targets. Log a once-per-process
  warning on a miss, and TODO lowering the floor once #4366 and
  workflow-server#1044 ship.
- Record workflow.run.spec_version / spec_version_source and the probe's
  latency or error on the start() span.
- Cap `wf inspect` replay's probed spec version at the CLI World's.
- Tests: cold spec-3 target, end-to-end probe timeout, 'latest' to another
  deployment and to self, retention error naming the target, compression
  off below spec 5, explicit-version attribution.
- Docs and changeset updates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>

Copy link
Copy Markdown
Member Author

@kvnloo good catch, you're right. A healthy stable target that missed the 2 s probe got stamped 6 and was rejected on pickup, the same outcome as main. The description's "errs low, never high" was wrong for older-major targets. Fixed in 764eafd:

  • The probe budget is now 10 s, the same budget wf inspect replay uses. healthCheck() returns on the first answer, so warm targets still cost one round-trip; only a real miss waits the full budget. Your scenario is now a regression test: a spec-3 target that answers after 5 s is stamped 3. That test fails with the old 2 s budget.
  • I kept the miss floor at 6, not 3, on purpose. On the Vercel World, spec 6 (slot identity) is the lowest version a v5 runtime can execute: it requires slot event ids, and runs below 6 get ULID ids. So a floor of 3 would fix unanswered stable targets but break every unanswered same-major target, which is the common case. No single fallback is safe for an unknown target, so I didn't add a test asserting one. Instead, a miss is now visible: the start() span records workflow.run.spec_version_source: probe-miss and the probe error, and it logs a warning once per process. There's a TODO to drop the floor to 3 once [world-vercel] Attest the executor's spec version on run_started #4366 and vercel/workflow-server#1044 let a v5 executor re-key an under-stamped fresh run.

I've corrected the description and the start.mdx bullet to match.

Copy link
Copy Markdown
Member Author

@pranaygp, going through your review summary:

  1. The 10 s budget is now paid once per deployment: there's a per-deployment cache, and 2 s after a miss (inline). The switch to 2 s with a floor of 3 is tracked in Lower the cross-deployment probe-miss floor to spec 3 once executors raise under-stamped runs #4401.
  2. A malformed JSON reply → 3 (inline).
  3. The branch is merged with main (spec 8) via update-branch. On the merged tree: packages/core src/ 2703 passed, the CLI inspect tests 120 passed. CI will run on the merged head.
  4. Added new 8/8 starter -> 7/7 executor: stamps 7, which checks the stamp in run_created, runInput and the queue option.
  5. Updated alternatives item 1 in the description: raises go to the executor's full version, structural crossings need a fresh log, and capability-only raises apply any time.
  6. Cookbook copy is written as shipped (inline).
  7. wf inspect now caps the source run's version too, not just the probe's.

Apply the same `>= 1` integer check `start()` uses, so a JSON reply with
`specVersion: 0` cannot stamp the replay with version 0. Also spell out, on
the probe-miss floor, the conditions the formal model (#4406, C3) found for
lowering it (#4401).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Peter Wielander <29887157+VaguelySerious@users.noreply.github.com>

@alangenfeld alangenfeld left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. One thing to check before merging: all 7 createHook({ experimental_force }) e2e tests failed on Vercel prod (quickjs), and "a run can take over its own earlier hook" timed out on local sveltekit. Neither failure shows up on main or #4366. The prod errors are UND_ERR_REQ_RETRY on the queue send, so this is probably flakiness, but the PR changes the specVersion that force-claim gates on. Worth a green re-run first.


Local agent review (`vercel-ai-gateway/anthropic/claude-opus-5.5`)

* target a caller can reach attests its version (a published beta that does
* not would be stamped below slot identity and fail). Even then a floor
* below `SPEC_VERSION_SUPPORTS_ATTRIBUTES` would refuse `attributes` on every
* miss. Tracked in vercel/workflow#4401.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking: lowering this floor to 3 depends on the run-spec-version-upgrade flag in vercel/workflow-server#1044 being on everywhere, not just on #1044 being merged. The flag is default-off there. Could this note (or #4401) say so, so nobody lowers the floor while the server still ignores executorSpecVersion?


Local agent review (`vercel-ai-gateway/anthropic/claude-opus-5.5`)

@pranaygp pranaygp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the quick round. Everything from the earlier review is addressed. Two things are left before I'd approve:

  1. A repeated miss keeps the deployment on the 2 s budget (inline). This reopens kvnloo's case for a stable target that answers slowly.
  2. Bundle Size Report fails because of this PR. The flow bundle grew about 55 KB raw (+13.4 KiB gzip) for hono and nextjs-turbopack against the ff80d62 baseline, which is over the 51,200-byte gate. About half the added source is comments. Please either trim it or apply allow-bundle-size-growth deliberately.

Missing tests for the cache:

  • the per-run public key is stripped from the cached entry;
  • a cache hit falls back to getEncryptionKeyForRun;
  • TTL expiry and eviction;
  • the 2 s budget is actually applied (the answer-after-miss test only advances 1 s).

Worth a sentence in the description: on a cache hit the caller fetches the run's symmetric key, where a probe would have let it seal to the run's public key. That trades an extra lookup and wider key access for skipping one queue round trip.

CI, other failures:

  • Python Conformance and python-node: also cancelled on every recent main run.
  • Local Dev nextjs-turbopack / sveltekit (stable): most likely flakes (an HMR test, and a force-claim test on the same-deployment path that never probes).
  • Vercel Prod nextjs-turbopack quickjs: a queues-proxy UND_ERR_REQ_RETRY.

None of these look caused by this PR; they need a rerun.

Once (1) and (2) are done this looks good to merge ahead of #1044. The over-stamp is live on main (spec 8), and this PR fixes it whenever the probe answers.

Comment thread packages/core/src/runtime/start.ts Outdated
const answered = probe?.format !== undefined;
byDeployment.delete(key);
byDeployment.set(key, {
at: Date.now(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every repeated miss rewrites at, so while starts to this deployment keep arriving within the 10-minute TTL, the entry stays fresh and the deployment never gets the 10 s budget back. A late answer is dropped instead of cached. The result: a stable (spec 3) target that takes more than 2 s to answer when cold is stamped 6 on every start, and it rejects those runs. That's kvnloo's case again. Suggest keeping the original miss timestamp when the entry was already a miss (so it expires and the next probe gets 10 s again), and/or caching an answer that arrives after the budget.

@pranaygp

Copy link
Copy Markdown
Contributor

https://github.com/vercel/vqs-server/pull/770

@alangenfeld @VaguelySerious I wonder if this PR is related to the failing CI tests?

Comment thread .changeset/cross-deployment-spec-version.md Outdated
VaguelySerious and others added 2 commits September 25, 2026 11:10
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Every miss rewrote the cache entry's timestamp, so while starts to a
deployment kept arriving it never got the full probe budget back, and a
target that answers only after the 2 s retry budget (a cold stable
target) was stamped with the spec-6 fallback on every start. A miss now
keeps a deployment on the short budget for 60 s from the first miss.

Also test that a cached answer looks the run key up rather than reusing
the probed one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com>

@pranaygp pranaygp left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The remaining blocker, a repeated probe miss pinning a deployment to the 2 s budget, is fixed in 29197a10f, along with a regression test and a check that a cache hit looks up the run key. We're accepting the flow-bundle growth (+55 KB raw) for this correctness fix, since the spec-8 over-stamp is live on main, so I added allow-bundle-size-growth. The Local Dev nextjs-webpack canary failure is the HMR suite, which doesn't go through start(). It needs a rerun (the bundle check also needs one to pick up the label).

Non-blocking follow-ups:

  • tests for cache TTL expiry and eviction;
  • a sentence in the description on the key-access trade-off: on a cache hit the caller fetches the run's symmetric key instead of sealing to its public key.

@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for c0aad6d (AI decision).

This is a genuine hang fix on main, but the defect it fixes is a main-only one: it is built entirely on spec versions and gates that do not exist on stable, whose @workflow/world tops out at SPEC_VERSION_CURRENT = 3 with no SPEC_VERSION_SUPPORTS_SLOT_IDENTITY/_ATTRIBUTES/_SEALED_LOG/_MAX_SUPPORTED, and whose start() has no attributes, experimental_retention, probe-supplied run key, or hookResumeInputVersion for the new resolver to touch. A stable caller already stamps the lowest version any executor accepts, so the over-stamp that hangs runs is not reachable there, and porting resolveCrossDeploymentSpecVersion would be a rewrite rather than a cherry-pick. The commit also carries additive surface (a per-deployment probe cache, a raised 10s probe budget, and five new telemetry attributes) and docs edits confined to docs/content/docs/v5/, which describe behavior stable does not have.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

c0aad6d01b65921a089174ded374bd6a2b515c78

This branch was successfully deployed

18 active deployments
Preview – workflow-swc-playground — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workflow-docs — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-vite-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-astro-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-express-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-hono-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – example-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workflow-tarballs — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workflow-web — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-python-workflow — 29197a10 Deployed Sep 25, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growth Accept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cross-deployment start() stamps the starter's spec version, not the executing deployment's

4 participants