Skip to content

fix: stop stamping spec 8 into runs an older runtime executes - #4415

Merged
alangenfeld merged 6 commits into
mainfrom
alangenfeld/e2e-target-spec-version
Sep 25, 2026
Merged

alangenfeld merged 6 commits into
mainfrom
alangenfeld/e2e-target-spec-version

Conversation

@alangenfeld

@alangenfeld alangenfeld commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary & Motivation

Both Python lanes (E2E Vercel Prod Tests (python - node) and E2E Python Conformance) have timed out on every run since #4193 raised the spec version to 8, which fails E2E Required Check on every PR. The deployed Python SDK writes spec 2 and reads at most 7, and two TypeScript writes stamp 8 into runs it executes:

  1. Runs. The harness starts runs as the deployment under test (VERCEL_DEPLOYMENT_ID is the target), so start() treats each as a same-deployment start and stamps this SDK's version. The deployment rejects every run and it stays pending:

    Vercel Queue push callback rejected: 1 validation error for RunStartedEvent
    spec_version
      Input should be less than or equal to 7 [type=less_than_equal, input_value=8, input_type=int]
    

    e2e-conformance.json already declares what a non-JS app's runtime supports, so it now also carries maxSpecVersion: the highest spec version the runtime accepts (7 for the pinned Python SDK). The harness stamps runs with the lower of that and this SDK's version. The health check cannot supply this number, since it reports the version a runtime writes (2 for Python), and stamping 2 disables everything gated at 4 or later, such as initial attributes. fix(core): stamp cross-deployment runs with the target deployment's spec version #4327 fixes the same class for real cross-deployment starts but cannot reach the harness, which never looks cross-deployment.

  2. Events written to someone else's run. resumeHook() stamped hook_received with SPEC_VERSION_CURRENT regardless of the run, so the run's runtime rejected its whole event log on replay (data.3.hook_received.specVersion: Input should be less than or equal to 7). This is not specific to the harness: any 5.x caller resuming a hook on a run executed by an older or non-JS runtime breaks that run the same way. Run.cancel() (run_cancelled) and the deployment guard's DEPLOYMENT_MISMATCH run_failed had the same over-stamp; cancelRun() already used the run's version, so the two cancel paths disagreed. All three now go through specVersionForRunWrite(), which stamps the lower of the run's version and this SDK's. Run.cancel() reads the run to learn its version. cancelRun(), wakeUpRun(), and reenqueueRun() in runs.ts already used the run's version but uncapped, so a run started by a newer SDK got a version this one cannot write; they now use the same helper, keeping their legacy fallback for a run with no recorded version.

Test Plan

  • Unit tests: specVersionForRunWrite() itself; hook_received (resume-hook.durable.test.ts), run_cancelled (run-cancel.test.ts), and the mismatch run_failed (deployment-guard.test.ts) are stamped with an older run's version, and cancelRun/wakeUpRun/reenqueueRun cap a newer run's at this SDK's while still treating an unversioned run as legacy. Each fails without its fix. The full @workflow/core unit suite passes.
  • Ran the suite locally against a workbench-python-workflow preview with CI's environment. Without the changes, warmup never picked up a run. With both, promiseAllWorkflow, the hook tests (hookWorkflow, hookWorkflow is not resumable via public webhook endpoint, hookCleanupTestWorkflow, hookWithSleepWorkflow), the three setAttributes > start: tests, and resilient start: addTenWorkflow completes when run_created returns 500 pass. Those last four passed at 7 before feat(hooks): createHook({ experimental_force: true }) takes a held token over #4193.

@changeset-bot

changeset-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 97172e5

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
Name Type
@workflow/core Patch
@workflow/builders Patch
@workflow/cli Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
example-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workflow-docs Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC
workflow-web Ready Ready Preview, v0 Sep 25, 2026 7:38pm UTC

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 97172e5 · Fri, 25 Sep 2026 20:01:47 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 403 (-45%) 💚 2449 🔴 (+124%) 🔻 2507 🔴 (+125%) 🔻 3007 🔴 (+150%) 🔻 30
TTFS stream 266 (+21%) 🔻 2333 🔴 (+122%) 🔻 2376 🔴 (+117%) 🔻 2768 🔴 (+94%) 🔻 30
TTFS hook + stream 562 (-36%) 💚 2760 🔴 (+108%) 🔻 2873 🔴 (+105%) 🔻 7040 🔴 (+351%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 694 (-20%) 💚 959 (-44%) 💚 2796 (+57%) 🔻 2842 (+52%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 3178 (+39%) 🔻 5489 (+51%) 🔻 6132 (+62%) 🔻 9831 (+36%) 🔻 10
STSO 1020 steps (inline) 126 (+26%) 🔻 158 (+22%) 🔻 175 (+14%) 257 (+27%) 🔻 1019
WO 1020 steps 155214 (+18%) 🔻 155214 (+18%) 🔻 155214 (+18%) 🔻 155214 (+18%) 🔻 1
CRTT first chunk (pooled) 75 (+4.2%) 116 (-18%) 💚 138 (-35%) 💚 148 (-58%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 99 (+1%) 159 (-17%) 336 (+11%) 719 (+46%) 256 (+14%) 10
size sweep (100/s, 160B-12KB) 107 (+3%) 154 (-32%) 197 (-44%) 399 (-37%) 133 (-37%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 96 (-30%) 146 (-23%) 191 (-41%) 708 (+23%) 397 (-9%) 3
replay eve-gpt-5.6-sol-2000t (1x) 120 (+15%) 172 (-13%) 225 (-26%) 340 (-38%) 244 (-46%) 2
replay eve-gpt-5.6-sol-2000t (2x) 109 (-23%) 207 (-37%) 289 (-41%) 709 (-25%) 367 (-22%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 131526ms → this run 154890ms (Δ +23364ms, +18%)

100-150 ms  ████████████████┃███████  main 899  this 649  -250
150-200 ms  ███░░░░░┃                 main 108  this 333  +225
200-250 ms  ┃                         main   8  this  24   +16
250-300 ms  ┃                         main   2  this   9    +7
300-350 ms  ┃                         main   0  this   1    +1
350-400 ms  ┃                         main   1  this   2    +1
400-450 ms  ┃                         main   1  this   0    -1
450-500 ms  ┃                         main   0  this   1    +1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▂█▁▁···     149 (-6%)   126 (-9%)  336 (+11%)  719 (+46%)  3000
sweep    ······▂█▁····  132.9 (-24%)  126 (-19%)  197 (-44%)  399 (-37%)  3000
gw 1x    ······▃█▁▁···  133.3 (-16%)  122 (-10%)  191 (-41%)  708 (+23%)  5295
eve 1x   ······▂█▂····  146.9 (-10%)   135 (±0%)  225 (-26%)  340 (-38%)  5186
eve 2x   ······▁█▃▁···  171.7 (-26%)  153 (-29%)  289 (-41%)  709 (-25%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▄▃▂▄██▄▁▆▆  122–176ms
sweep    ██▄▂▃▃▁▅▃▄  120–150ms
gw 1x    ▆█▄▂▂▂▁▁▃▁  115–179ms
eve 1x   ▅▁▂▃▂▃▇█▄▃  131–169ms
eve 2x   ▁▂▁▁▁▄▆▇█▅  140–224ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▁█▇▄▁▄▁  132–135ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▂▂▄▆▁▄▅▂█▅  37–53ms
sweep    ▁▆▄▁▅▂▄█▆▄  50–66ms
gw 1x    ███▁▁▄▄▃▃▂  34–47ms
eve 1x   ▆▁▃▂▃█▆▃▁▇  25–33ms
eve 2x   ▆▇▁▂▅▂▄▁▆█  22–34ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • a run can take over its own earlier hook (astro · local-dev / local / node / stable)
  • abortFromStepWorkflow: step abort cancels an in-flight sibling step (hono · vercel-prod / vercel / node / production)
  • health check (CLI) - workflow health command reports healthy endpoints (nextjs-webpack · local-dev / local / node / canary)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • cold-start-warmup · suite warmup (tanstack-start) · at 19:40:52Z · abandoned wrun_01M3D1AWCZS1S3Z1EZQTJGZ4J0
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 19:45:24Z · abandoned wrun_01M3D1K9CX4G70B4Q4DZP2X1EH

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3878 0 739 4617
✅ 💻 Local Development 4238 0 550 4788
✅ 📦 Local Production 4238 0 550 4788
✅ 🐘 Local Postgres 4238 0 550 4788
✅ 🪟 Windows 340 0 2 342
✅ 🌐 Cross-language Conformance 68 0 84 152
✅ vercel-http-transport 873 0 153 1026
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 591 0 93 684
Total 18491 0 2721 21212
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 141 0 30
✅ astro-quickjs 141 0 30
✅ example-node 141 0 30
✅ example-quickjs 141 0 30
✅ express-node 141 0 30
✅ express-quickjs 141 0 30
✅ fastify-node 141 0 30
✅ fastify-quickjs 141 0 30
✅ hono-node 141 0 30
✅ hono-quickjs 141 0 30
✅ nest-node 141 0 30
✅ nest-quickjs 141 0 30
✅ nextjs-turbopack-node 168 0 3
✅ nextjs-turbopack-quickjs 168 0 3
✅ nextjs-webpack-node 168 0 3
✅ nextjs-webpack-quickjs 168 0 3
✅ nitro-node 141 0 30
✅ nitro-quickjs 141 0 30
✅ nuxt-node 141 0 30
✅ nuxt-quickjs 141 0 30
✅ python-node 66 0 105
✅ sveltekit-node 160 0 11
✅ sveltekit-quickjs 160 0 11
✅ tanstack-start-node 141 0 30
✅ tanstack-start-quickjs 141 0 30
✅ vite-node 141 0 30
✅ vite-quickjs 141 0 30

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 142 0 29
✅ astro-stable-quickjs 142 0 29
✅ express-stable-node 142 0 29
✅ express-stable-quickjs 142 0 29
✅ fastify-stable-node 142 0 29
✅ fastify-stable-quickjs 142 0 29
✅ hono-stable-node 142 0 29
✅ hono-stable-quickjs 142 0 29
✅ nest-stable-node 142 0 29
✅ nest-stable-quickjs 142 0 29
✅ nextjs-turbopack-canary-node 170 0 1
✅ nextjs-turbopack-canary-quickjs 170 0 1
✅ nextjs-turbopack-stable-node 170 0 1
✅ nextjs-turbopack-stable-quickjs 170 0 1
✅ nextjs-webpack-canary-node 170 0 1
✅ nextjs-webpack-canary-quickjs 170 0 1
✅ nextjs-webpack-stable-node 170 0 1
✅ nextjs-webpack-stable-quickjs 170 0 1
✅ nitro-stable-node 142 0 29
✅ nitro-stable-quickjs 142 0 29
✅ nuxt-stable-node 142 0 29
✅ nuxt-stable-quickjs 142 0 29
✅ sveltekit-stable-node 161 0 10
✅ sveltekit-stable-quickjs 161 0 10
✅ tanstack-start-node 142 0 29
✅ tanstack-start-quickjs 142 0 29
✅ vite-stable-node 142 0 29
✅ vite-stable-quickjs 142 0 29

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 142 0 29
✅ astro-stable-quickjs 142 0 29
✅ express-stable-node 142 0 29
✅ express-stable-quickjs 142 0 29
✅ fastify-stable-node 142 0 29
✅ fastify-stable-quickjs 142 0 29
✅ hono-stable-node 142 0 29
✅ hono-stable-quickjs 142 0 29
✅ nest-stable-node 142 0 29
✅ nest-stable-quickjs 142 0 29
✅ nextjs-turbopack-canary-node 170 0 1
✅ nextjs-turbopack-canary-quickjs 170 0 1
✅ nextjs-turbopack-stable-node 170 0 1
✅ nextjs-turbopack-stable-quickjs 170 0 1
✅ nextjs-webpack-canary-node 170 0 1
✅ nextjs-webpack-canary-quickjs 170 0 1
✅ nextjs-webpack-stable-node 170 0 1
✅ nextjs-webpack-stable-quickjs 170 0 1
✅ nitro-stable-node 142 0 29
✅ nitro-stable-quickjs 142 0 29
✅ nuxt-stable-node 142 0 29
✅ nuxt-stable-quickjs 142 0 29
✅ sveltekit-stable-node 161 0 10
✅ sveltekit-stable-quickjs 161 0 10
✅ tanstack-start-node 142 0 29
✅ tanstack-start-quickjs 142 0 29
✅ vite-stable-node 142 0 29
✅ vite-stable-quickjs 142 0 29

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 142 0 29
✅ astro-stable-quickjs 142 0 29
✅ express-stable-node 142 0 29
✅ express-stable-quickjs 142 0 29
✅ fastify-stable-node 142 0 29
✅ fastify-stable-quickjs 142 0 29
✅ hono-stable-node 142 0 29
✅ hono-stable-quickjs 142 0 29
✅ nest-stable-node 142 0 29
✅ nest-stable-quickjs 142 0 29
✅ nextjs-turbopack-canary-node 170 0 1
✅ nextjs-turbopack-canary-quickjs 170 0 1
✅ nextjs-turbopack-stable-node 170 0 1
✅ nextjs-turbopack-stable-quickjs 170 0 1
✅ nextjs-webpack-canary-node 170 0 1
✅ nextjs-webpack-canary-quickjs 170 0 1
✅ nextjs-webpack-stable-node 170 0 1
✅ nextjs-webpack-stable-quickjs 170 0 1
✅ nitro-stable-node 142 0 29
✅ nitro-stable-quickjs 142 0 29
✅ nuxt-stable-node 142 0 29
✅ nuxt-stable-quickjs 142 0 29
✅ sveltekit-stable-node 161 0 10
✅ sveltekit-stable-quickjs 161 0 10
✅ tanstack-start-node 142 0 29
✅ tanstack-start-quickjs 142 0 29
✅ vite-stable-node 142 0 29
✅ vite-stable-quickjs 142 0 29

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 170 0 1
✅ nextjs-turbopack-quickjs 170 0 1

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 84

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 141 0 30
✅ express 141 0 30
✅ hono 141 0 30
✅ nextjs-turbopack 168 0 3
✅ nitro 141 0 30
✅ vite 141 0 30

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 141 0 30
✅ express 141 0 30
✅ nextjs-turbopack 168 0 3
✅ vite 141 0 30

📋 View full workflow run

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 42 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ step-vs-timer-early-settlement completed 8 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 265.3 KiB (±0) 96.1 KiB (-27 B) 1.93 MiB (+327 B)
nextjs-turbopack 272.8 KiB (±0) 426 B (±0) 922.8 KiB (+130 B)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

97172e5 · run

@alangenfeld alangenfeld changed the title test(e2e): stamp runs with a non-JS target's spec version fix: stop stamping spec 8 into runs an older runtime executes Sep 25, 2026
@alangenfeld
alangenfeld marked this pull request as ready for review September 25, 2026 17:17
@alangenfeld
alangenfeld requested a review from a team as a code owner September 25, 2026 17:17
The harness starts runs as the deployment under test (VERCEL_DEPLOYMENT_ID
is the target), so start() treats them as same-deployment and stamps this
SDK's spec version. Since #4193 that is 8, and the Python SDK rejects every
run above 7 ('1 validation error for RunStartedEvent spec_version: Input
should be less than or equal to 7'), leaving it pending. Both Python lanes
have timed out on every run since.

For apps that declare an e2e-conformance.json, probe the target once with
the health check and stamp the lower of its answer and this SDK's version.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
resumeHook() stamped hook_received with SPEC_VERSION_CURRENT whatever
the run's version. Since #4193 that is 8, and a runtime built against a
lower version rejects the event when it replays the run. The Python SDK
rejects the whole event log ('data.3.hook_received.specVersion: Input
should be less than or equal to 7'), so every hook test on the Python
lanes timed out even once runs were stamped for the target.

Stamp the lower of the run's version and this SDK's. The payload is
already encoded for the run's version.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
The health check reports the spec version a runtime writes, not the
highest it reads. For the Python SDK those are 2 and 7, and stamping 2
disables everything gated at 4 or later: start() rejects initial
attributes, and the resilient-start path changes. Four conformance
tests that passed at 7 before #4193 failed.

e2e-conformance.json already declares what an app's runtime supports,
so it now carries maxSpecVersion (7 for the pinned Python SDK) and the
harness stamps the lower of that and this SDK's version.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
A maxSpecVersion above what the app's runtime accepts leaves every run
pending, which reads as a stalled queue. Say so where the warmup stall
is reported, and name the exact constant and lockfile to keep it in
step with.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
… version

Run.cancel() and the DEPLOYMENT_MISMATCH run_failed write to a run from a
process that is often not its executor, and stamped SPEC_VERSION_CURRENT
like hook_received did. A runtime that validates the whole log (the
Python SDK) rejects it. cancelRun() already used the run's version, so
the two cancel paths disagreed.

Move the rule into specVersionForRunWrite() and use it for all three
writes. Run.cancel() now reads the run to learn its version.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@alangenfeld
alangenfeld force-pushed the alangenfeld/e2e-target-spec-version branch from 9e8bb00 to 176fbd4 Compare September 25, 2026 18:57

@VaguelySerious VaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI review: no blocking issues

'use step';
const world = await this.#lazyWorldPromise;
// The caller is often not the run's executor, so stamp the run's version.
const run = await world.runs.get(this.runId, { resolveData: 'none' });

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Nit

This adds a runs.get round trip to every external Run.cancel(), before the write. It's correct, and cancelRun() in runs.ts already pays the same read, so both cancel paths are now consistent. Cancel isn't a hot path, so this is only worth a mention. I checked it end to end against world-local with a scratch test: on a run at SPEC_VERSION_CURRENT - 1, run_cancelled is stored at that version, the run ends cancelled, and cancelling a missing run still fails with a not-found error.

* event log over a single such event. A run with no recorded version gets this
* SDK's, as before.
*/
export function specVersionForRunWrite(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Nit

cancelRun(), wakeUpRun() and reenqueueRun() in runtime/runs.ts still stamp run.specVersion directly, uncapped. For a run newer than this SDK, they stamp a version this SDK can't produce, which this helper caps. Routing them through specVersionForRunWrite would give all external writes one rule. It's harmless for the reader, so a follow-up is fine.

…rsion

They stamped run.specVersion directly, so a run started by a newer SDK
got a version this one cannot write. Route them through
specVersionForRunWrite(), which now takes the fallback for a run with no
recorded version, so these keep treating such a run as legacy.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

@TooTallNate TooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Approved — no blocking issues found in 97172e5.

Reviewed the target-spec E2E wrapper, cross-runtime event stamping, and the latest cancelRun/wakeUpRun/reenqueueRun changes, including preservation of the legacy fallback. The earlier review feedback is addressed.

Validation: all 107 tests passed across run-spec-version, run-cancel, runs, resume-hook.durable, deployment-guard, and e2e/utils. Both Python CI lanes and the E2E Required Check also pass.

The remaining red Tarballs Preview Smoke Checks job was cancelled while waiting for the preview deployment; its endpoint-verification step never ran. That check does not provide evidence of a code regression in this diff.

@alangenfeld
alangenfeld merged commit 20aa656 into main Sep 25, 2026
184 of 185 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-target-spec-version branch September 25, 2026 21:17
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 20aa656 (AI decision).

This fixes a breakage that only exists on main: spec version 8 (from #4193) being stamped into runs executed by the pinned Python SDK, which reads at most 7. stable is still at SPEC_VERSION_CURRENT = 3, has no workbench/python or e2e-conformance.example.json, no deployment-guard.ts, and no resumeContext.runSpecVersion on hooks (it reads hook.specVersion), so most of the commit targets code and CI lanes that do not exist there and the rest depends on main-only shape. The genuinely portable part is the correctness rule itself — stamping hook_received, run_cancelled, and the mismatch run_failed with min(run.specVersion, SDK version) via specVersionForRunWrite() — which a human could reimplement against stable's structures and force through if cross-SDK over-stamping is judged a real risk at spec 3.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

20aa6562f1f70917013ddad5716eb40661435fdc

This branch was successfully deployed

18 active (17 outdated) deployments
Preview – workflow-docs — 97172e54 Deployed Sep 25, 2026 by vercel[bot]
Preview – workflow-swc-playground — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-vite-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workflow-tarballs — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-hono-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-astro-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-express-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – example-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workflow-web — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Preview – workbench-python-workflow — 9e8bb00d Deployed Sep 25, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants