Skip to content

test(next): cover HMR rebuild convergence - #3471

Open
NathanColosimo wants to merge 12 commits into
codex/next-hmr/startup-watchfrom
codex/next-hmr/e2e
Open

NathanColosimo wants to merge 12 commits into
codex/next-hmr/startup-watchfrom
codex/next-hmr/e2e

Conversation

@NathanColosimo

@NathanColosimo NathanColosimo commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • make expected hot, full, and skipped rebuild counts explicit
  • verify body-only edits stay hot while import, add, and remove changes rediscover workflows
  • verify editing an imported dependency updates workflow execution on stable and canary Next.js
  • test Next.js 16.3.1 and 16.3.1-canary.21 across Turbopack, webpack, Node.js, and QuickJS

Testing

  • pnpm --filter @workflow/next build
  • pnpm --filter @workflow/next test
  • focused local Next.js 16.3.1-canary.21 HMR scenario with Turbopack and webpack
  • GitHub Actions Next.js local-development matrix

Stack

  1. fix(next): converge workflow HMR rebuilds #3333: rebuild convergence and unit coverage
  2. fix(next): watch before initial build #3470: close the startup observation gap
  3. This PR: stable/canary end-to-end HMR coverage

Stack created with GitHub Stacks CLI • Give Feedback 💬

@changeset-bot

changeset-bot Bot commented Aug 11, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: da9995c

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actions Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • sleepWinsRaceWorkflow (tanstack-start · local-dev / local / node)
  • stepWinsRaceWorkflow (vite · local-dev / local / node / stable)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • cold-start-warmup · suite warmup (astro) · at 22:49:32Z · abandoned wrun_01M2H1RAC17ZNRJTTXYA927CP0
  • cold-start-warmup · suite warmup (tanstack-start) · at 22:49:38Z · abandoned wrun_01M2H1RGMB682N0MFR14H529ET
  • run-pickup-stall · hookWorkflow (nuxt) · at 22:50:05Z · abandoned wrun_01M2H1SJEZEVAGHWBYZ4FFKXR0

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3662 0 685 4347
✅ 💻 Local Development 2718 0 502 3220
✅ 📦 Local Production 3998 0 510 4508
✅ 🐘 Local Postgres 3998 0 510 4508
✅ 🪟 Windows 320 0 2 322
✅ 🌐 Cross-language Conformance 68 0 74 142
✅ vercel-http-transport 823 0 143 966
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 557 0 87 644
Total 16171 0 2513 18684
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 133 0 28
✅ astro-quickjs 133 0 28
✅ example-node 133 0 28
✅ example-quickjs 133 0 28
✅ express-node 133 0 28
✅ express-quickjs 133 0 28
✅ fastify-node 133 0 28
✅ fastify-quickjs 133 0 28
✅ hono-node 133 0 28
✅ hono-quickjs 133 0 28
✅ nest-node 133 0 28
✅ nest-quickjs 133 0 28
✅ nextjs-turbopack-node 158 0 3
✅ nextjs-turbopack-quickjs 158 0 3
✅ nextjs-webpack-node 158 0 3
✅ nextjs-webpack-quickjs 158 0 3
✅ nitro-node 133 0 28
✅ nitro-quickjs 133 0 28
✅ nuxt-node 133 0 28
✅ nuxt-quickjs 133 0 28
✅ python-node 66 0 95
✅ sveltekit-node 152 0 9
✅ sveltekit-quickjs 152 0 9
✅ tanstack-start-node 133 0 28
✅ tanstack-start-quickjs 133 0 28
✅ vite-node 133 0 28
✅ vite-quickjs 133 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 160 0 1
✅ nextjs-turbopack-canary-quickjs 160 0 1
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 160 0 1
✅ nextjs-webpack-canary-quickjs 160 0 1
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 27
✅ astro-stable-quickjs 134 0 27
✅ express-stable-node 134 0 27
✅ express-stable-quickjs 134 0 27
✅ fastify-stable-node 134 0 27
✅ fastify-stable-quickjs 134 0 27
✅ hono-stable-node 134 0 27
✅ hono-stable-quickjs 134 0 27
✅ nest-stable-node 134 0 27
✅ nest-stable-quickjs 134 0 27
✅ nextjs-turbopack-canary-node 160 0 1
✅ nextjs-turbopack-canary-quickjs 160 0 1
✅ nextjs-turbopack-stable-node 160 0 1
✅ nextjs-turbopack-stable-quickjs 160 0 1
✅ nextjs-webpack-canary-node 160 0 1
✅ nextjs-webpack-canary-quickjs 160 0 1
✅ nextjs-webpack-stable-node 160 0 1
✅ nextjs-webpack-stable-quickjs 160 0 1
✅ nitro-stable-node 134 0 27
✅ nitro-stable-quickjs 134 0 27
✅ nuxt-stable-node 134 0 27
✅ nuxt-stable-quickjs 134 0 27
✅ sveltekit-stable-node 153 0 8
✅ sveltekit-stable-quickjs 153 0 8
✅ tanstack-start-node 134 0 27
✅ tanstack-start-quickjs 134 0 27
✅ vite-stable-node 134 0 27
✅ vite-stable-quickjs 134 0 27

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 160 0 1
✅ nextjs-turbopack-quickjs 160 0 1

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 74

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 133 0 28
✅ express 133 0 28
✅ hono 133 0 28
✅ nextjs-turbopack 158 0 3
✅ nitro 133 0 28
✅ vite 133 0 28

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 133 0 28
✅ express 133 0 28
✅ nextjs-turbopack 158 0 3
✅ vite 133 0 28

📋 View full workflow run

@vercel

vercel Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
example-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workflow-docs Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC
workflow-web Ready Ready Preview, v0 Sep 14, 2026 10:49pm UTC

@github-actions

github-actions Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit da9995c · Mon, 14 Sep 2026 23:03:56 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1565 (+18%) 🔻 1657 🔴 (+15%) 1832 🔴 (+25%) 🔻 2114 🔴 (+31%) 🔻 30
TTFS stream 1547 (+706%) 🔻 1645 🔴 (+15%) 🔻 1665 🔴 (+14%) 1681 🔴 (+12%) 30
TTFS hook + stream 549 (-4.4%) 2044 🔴 (+11%) 2125 🔴 (+8.4%) 6287 🔴 (+173%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 468 (+3.3%) 684 (+1.2%) 704 (-62%) 💚 6658 (+259%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 1583 (-5.7%) 3762 (+15%) 🔻 4922 (+13%) 10307 (+30%) 🔻 10
STSO 1020 steps (inline) 137 (+8.7%) 162 (+1.9%) 178 (-1.1%) 261 (+23%) 🔻 1019
WO 1020 steps 163058 (+2.6%) 163058 (+2.6%) 163058 (+2.6%) 163058 (+2.6%) 1
CRTT first chunk (pooled) 60 (-34%) 💚 86 (-42%) 💚 96 (-53%) 💚 161 (-27%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 73 (-49%) 169 (-58%) 329 (-44%) 676 (-18%) 165 (-42%) 10
size sweep (100/s, 160B-12KB) 81 (-39%) 144 (-64%) 221 (-58%) 353 (-62%) 128 (-51%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 89 (-21%) 160 (-40%) 265 (-33%) 644 (+8%) 324 (-20%) 3
replay eve-gpt-5.6-sol-2000t (1x) 81.5 (-50%) 160 (-38%) 247 (-27%) 567 (-25%) 359 (-48%) 2
replay eve-gpt-5.6-sol-2000t (2x) 72 (-43%) 188 (-59%) 265 (-59%) 465 (-62%) 254 (-41%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 158739ms → this run 162934ms (Δ +4195ms, +3%)

100-150 ms  ███████┃█████████         main 512  this 243  -269
150-200 ms  ████████████████░░░░░░░┃  main 477  this 735  +258
200-250 ms  ┃                         main  26  this  29    +3
250-300 ms  ┃                         main   2  this   9    +7
300-350 ms  ┃                         main   0  this   3    +3
400-450 ms  ┃                         main   2  this   0    -2
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▃█▂▁···  142.8 (-47%)  126 (-43%)  329 (-44%)  676 (-18%)  3000
sweep    ······▃█▁····  125.8 (-56%)  120 (-48%)  221 (-58%)  353 (-62%)  3000
gw 1x    ·····▁▄█▂▁···  138.9 (-31%)  122 (-31%)  265 (-33%)   644 (+8%)  5295
eve 1x   ·····▁▄█▂▁···  132.9 (-37%)  114 (-36%)  247 (-27%)  567 (-25%)  5186
eve 2x   ·····▁▂█▃▁···  159.2 (-51%)  145 (-54%)  265 (-59%)  465 (-62%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▃▇▄▄█▄▃▅▃▁  114–178ms
sweep    █▂▂▆▂▂▄▅▁▁  116–146ms
gw 1x    █▄▂█▄▁▄▂▃▁  117–173ms
eve 1x   ▂▁▂▄▂▂▅▅█▂  108–183ms
eve 2x   ▃▂█▅▁▃▅█▇▁  125–197ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▆▅▄▄▁▃█  124–128ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▄▇▃█▇▃▅▄▃▁  33–53ms
sweep    ▁▁▆█▄▆▆█▄▃  46–64ms
gw 1x    █▃▁▃▄▂▄▂▃▂  35–58ms
eve 1x   ▃▂▁▄▃▁█▇▄▅  21–32ms
eve 2x   ██▅▄▅▃▄▁▄▅  21–31ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). █ = main, ┃ = this run, ░ = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

Comment thread packages/core/e2e/dev.test.ts Outdated
@github-actions

github-actions Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 250.7 KiB (±0) 93.0 KiB (±0) 1.89 MiB (±0)
nextjs-turbopack 257.2 KiB (±0) 426 B (±0) 897.9 KiB (±0)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

da9995c · run

@VaguelySerious

Copy link
Copy Markdown
Member

AI Review: Blocking

The Local Dev lanes are red for a specific, reproducible reason, and it is worth deciding what it means before this lands.

8 of 8 Next dev lanes fail identically (turbopack + webpack × stable + canary × node + quickjs), all on the same assertion, freshly as of 22:53 today:

FAIL packages/core/e2e/dev.test.ts > should follow Next flow-route HMR rebuild rules for body-only changes
Timed out after 90000ms waiting for step HMR update to affect workflow execution.
Last error: Workflow run "wrun_..." failed: Step "step//./workflows/hmr-fuzz-serde//hmrFuzzSerdeStep"
failed after 3 retries: Class "HmrFuzzBox" not found. Make sure the class is registered with
registerSerializationClass.

Identical failure across all eight is not flakiness.

What changed. The test name predates this PR, but the assertion does not. This PR adds hmrFuzzSerdeStep(value: HmrFuzzBox) to the hmr-fuzz-serde fixture and has the workflow call await hmrFuzzSerdeStep(new HmrFuzzBox(workflowValue)). That makes a custom-serde class cross the step boundary as an argument for the first time in this suite, so the step has to deserialize it, which requires the class in the step bundle's registry. It isn't there.

Why I think this is a product gap rather than a bad fixture, with the caveat that I did not reproduce it locally:

  • The fixture uses the declarative convention (static classId plus the workflow-serialize / workflow-deserialize symbols), which is the same convention workbench/example/workflows/serde-models.ts uses.
  • Hybrid modules that define both serializers and steps do work in static builds: workbench/example/workflows/99_e2e.ts has 94 'use step' occurrences alongside 3 serializer classes and passes the Vercel Prod lanes.
  • The fixture's distinguishing property is that the class and the step live in the same file and that file is being rewritten by HMR. So the delta against the known-working cases is the dev rebuild path, not the convention.

If that is right, the registration gap is a bug in the dev builder for hybrid serde+step modules, and fixing it is out of scope for a PR whose job is test coverage. The stack under this one (#3333, #3470) does not address it. This is the same family of problem as #3639, which exists because hybrid serializer sources need special handling at bundle boundaries.

Recommended: triage the failure first. If it is a product bug, split it out and let this PR land the coverage that passes; if the fixture simply needs an explicit registerSerializationClass, that is a two-line change here.

Not blocking, for the record: the Unit Tests (windows-latest) failure is unrelated and is the known world-local hook-convergence flake (src/storage.test.ts > converges same-hook creation across separate storage instances to one event, expected 25 got 26). Windows unit passed on the two most recent main runs, so a re-run should clear it.

DCO is red for the org-membership reason explained on #3333, not for anything in this diff.

@VaguelySerious VaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI review: blocking issues found

value: `serde-body-${iteration}`,
}) satisfies ExpectedWorkflowResult,
source: (iteration: number) => `export class HmrFuzzBox {
static classId = 'HmrFuzzBox';

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Blocking

All 8 E2E Local Dev Tests lanes on this PR fail deterministically (run 34905603125, head da9995c5):

Step "step//./workflows/hmr-fuzz-serde//hmrFuzzSerdeStep" failed after 3 retries:
Class "HmrFuzzBox" not found. Make sure the class is registered with registerSerializationClass.

Cause is this static classId. The SWC plugin emits a registration IIFE next to every class carrying the serialize/deserialize symbols:

__wf_reg.set(__wf_id, __wf_cls);
if (!Object.prototype.hasOwnProperty.call(__wf_cls, "classId"))
  Object.defineProperty(__wf_cls, "classId", { value: __wf_id, ... });

The .set() is unconditional under the path-derived id, but the defineProperty is guarded by hasOwnProperty. A class that declares its own static classId keeps it. So the registry holds only class//./workflows/hmr-fuzz-serde//HmrFuzzBox while HmrFuzzBox.classId stays 'HmrFuzzBox'. The serialize reducer stamps constructor.classId into the payload, and deserialize looks that exact string up, finds nothing, and throws the message above.

I drove the real prebuilt plugin over this fixture and ran the runtime's reducer pair against the output, with and without the static, in both plugin modes:

fixture mode result
with static classId step throws Class "HmrFuzzBox" not found
with static classId workflow throws Class "HmrFuzzBox" not found
without static classId step round-trips
without static classId workflow round-trips

Attribution, since it is not obvious from the diff: this line is inherited from #3470, where it is inert. On the base branch HmrFuzzBox is constructed and compared but never crosses a step boundary, so the reducer never runs and #3470 is E2E-green. This PR adds hmrFuzzSerdeStep (here and at line 1108) plus five await hmrFuzzSerdeStep(new HmrFuzzBox(workflowValue)) call sites, which is the first code to serialize the class. The latent poisoned id becomes fatal here.

Fix: drop static classId from both fixture copies (this one and the one at line 1095). That matches the working precedent in workbench/example/workflows/serde-models.ts, where Vector declares only the two symbol statics and lets the plugin assign the id.

kind: 'workflow',
value: `serde-body-${iteration}`,
}) satisfies ExpectedWorkflowResult,
source: (iteration: number) => `export class HmrFuzzBox {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Note

The HmrFuzzBox fixture source is written out verbatim at two sites in this file (here and around line 1094), and the hmrFuzzSerdeStep call appears at five places. When the blocking fix above lands, all copies have to change together or half the matrix keeps failing. A single shared template string for the serde fixture would make that a one-line edit and remove the class of bug where the two copies drift.

@VaguelySerious

Copy link
Copy Markdown
Member

AI Review: Note

Separate from the review above, on the other red check: Unit Tests (windows-latest) (job 104181466190) fails one test out of 19 files, and it is in a package this stack does not touch.

FAIL  src/storage.test.ts > Storage > concurrent entity-creation races > converges same-hook creation across separate storage instances to one event
AssertionError: expected [ { …(7) }, { …(7) }, { …(7) }, …(23) ] to have a length of 25 but got 26
 ❯ src/storage.test.ts:4444:27

Why I read this as pre-existing rather than caused by the stack:

  • The three-PR stack's diff against main is 23 files, none under packages/world-local.
  • git diff --stat origin/main <head> -- packages/world-local packages/world/src is empty, so the package is byte-identical to main at this head.
  • No world-local commits landed on main since the merge-base, so the branch is not missing a fix.
  • Unit Tests (windows-latest) was green on the last eight main runs of tests.yml, and has never failed on any prior run of the three stack branches.
  • The test itself dates to fix(world-local,world-postgres): make duplicate hook_created idempotent #2295 (2026-06-11).

The symptom is worth recording for whoever owns the flake: every one of the 25 iterations passed its per-iteration assertions (exactly one fulfilled, exactly one rejected with EntityConflictError), yet 26 hook_created events were persisted. A loser rejected and still published its event. That is the duplicate-publication case the test was written to catch, so it looks like a underlying race in the cross-instance convergence path rather than a bad assertion.

Caveat on the evidence: eight green main runs bounds nothing about the rate, and I did not scan all of history for this test name. The load-bearing part is non-causality (identical code, zero touched files), not the green sample.

This branch was successfully deployed

18 active deployments
Preview – workflow-swc-playground — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workflow-docs — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-astro-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-vite-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – example-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-express-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workflow-tarballs — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-hono-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workflow-web — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Preview – workbench-python-workflow — da9995c5 Deployed Sep 14, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants