[DO NOT MERGE] CI dry-run: QuickJS as the default workflow VM engine - #3253
TooTallNate wants to merge 36 commits into
Conversation
🦋 Changeset detectedLatest commit: 8a78801 The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (1 failed)express-quickjs (1 failed):
📦 Local Production (1 failed)astro-stable-node (1 failed):
E2E Test SummarySummary
Details by Category❌ ▲ Vercel Production
✅ 💻 Local Development
❌ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ vercel-multi-region
|
VaguelySerious
left a comment
There was a problem hiding this comment.
AI review: blocking issues found
…KFLOW_SNAPSHOT_THRESHOLD)
…based PRNG fast-forward, unified host-callback list, lifecycle hardening - SnapshotMetadata gains eventCount, rngDraws and formatVersion. The max-events guard now compares restored total + delta (both at entry and per loop turn) — previously a run that kept snapshotting could never accumulate enough delta to trip the ceiling it exists for. - Correlation-id generation is position-based across snapshots: the runtime seeds from the BASE seed and fast-forwards the persisted draw count instead of mixing the snapshot cursor into the seed. Ids are now identical across snapshot generations AND identical to a no-snapshot run, so overlapping invocations straddling a snapshot save still collide on the world's dedup (new test pins restored ids == full-replay ids). Snapshots without a draw count fall back to full replay. - Host callbacks are declared in ONE list that drives both the fresh-boot install and the restore re-registration, so adding a callback can't silently skip the restore path. - Preloaded events are used again with snapshotting enabled (the first qualifying suspension skips its save — no cursor yet); short runs keep the zero-round-trip fast path. - Snapshot persist runs off the response path (waitUntil), with a 32MB plaintext size ceiling (skip + warn). Loads that fail format/shape checks warn instead of silently miming a miss. Terminal deletes are gated on a snapshot actually existing and now also fire on the runGone path; a server-side TTL remains the backstop for unobserved cancellations.
…ck into node:vm) - useQuickJSVm defaults to the QuickJS engine when neither the run's stamped executionContext.workflowVm nor WORKFLOW_VM specifies one; WORKFLOW_VM=node is the explicit node:vm opt-in. - CI matrix inverted to match: quickjs legs leave WORKFLOW_VM unset so the default-selection path is exercised end to end; node legs opt in explicitly (labels/artifacts unchanged via MATRIX_VM). - Entrypoint tests that assert node replay-loop internals against mock worlds pin WORKFLOW_VM=node (the quickjs path would instantiate a WASM VM per call). - Docs + changeset updated. Runs keep the engine stamped at start().
1ff3cc8 to
6af2c2c
Compare
📊 Workflow Benchmarkscommit Backend:
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 170410ms → this run 864672ms (Δ +694262ms, +407%) 1020 steps (queue-hop) Cumulative STSO time: 9310ms over 2 samples No 📜 Previous results (4)c671016Mon, 10 Aug 2026 20:52:43 GMT · run logs
d41872aMon, 10 Aug 2026 19:58:41 GMT · run logs
74bd053Tue, 04 Aug 2026 01:01:14 GMT · run logs
6af2c2cFri, 31 Jul 2026 23:40:21 GMT · run logs
ℹ️ Metric definitions & methodologyThe collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor ( Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the |
|
Ran the event-log-race-repro storm locally against this branch on world-local and world-postgres (24 step-storm + 24 hook-storm per condition, concurrency 6,
¹ partial runs (21 and 19 landed of 48) — the launch budget cut them short. Caveat before reading the zero: under QuickJS not a single storm run completed — every run blew past the harness's 240s cap ( A fair comparison needs |
…hreads Merge resolution — main's #3048 finals carried into the inline-loop architecture: - namespace + run-origin nextTraceCarrier threaded through runWorkflowWithQuickJS into every publish (step handoffs, hook_conflict requeue, wait continuations, immediate requeues) - suspended-exit requeues converted to FRESH messages (never { timeoutSeconds } visibility-redelivery of the current message — the hookInput redelivery trap fixed on #3048); exit wait sweep enqueues the continuation for the soonest unscheduled wait directly - entrypoint-side hookInput materialization dropped in favor of main's engine-agnostic prologue re-ensure in runtime.ts (with #3230's (runId, resumeId) claim protocol); dispatch stays inside the replay loop's try so engine failures classify into run_failed - interrupt handler keeps the perf branch's per-burst mutable budget, with main's configurable getReplayTimeoutMs() as the ceiling Review fixes (PR #3049 threads): - CRITICAL overflow wedge: overflow steps are handed to the queue in the same turn their step_created is written, BEFORE the event feed — the feed always observes those writes and continued the loop, so the old handoff was unreachable on the only turn that classified the steps as fresh (the cause of promiseRaceStressTestWorkflow hanging in the quickjs CI legs) - backstop gating: the deliveryAttempt > 1 gate (common case on worlds that advance attempts on routine redeliveries) is replaced with the node engine's ownership decision table — lease-active steps owned by another message arm a DELAYED backstop for the lease remainder under an epoch-scoped key; owner redeliveries and expired/unstamped steps dispatch immediately under the bare-correlationId key. Ownership is derived host-side from observed step_started/step_retrying events - ack-without-requeue: inline step terminals the feed has not surfaced raise the requeue signal, so the loop never acks with durably written terminals and nothing scheduled to consume them - idempotency keys bucketed by purpose (dispatch / backstop:<epoch> / retry:<n>) so worlds that retire used keys cannot swallow a later publish for the same step - live-feed terminal buffering: step/wait/attr terminals arriving before this VM constructs the corresponding resolver are buffered (__terminalBuffer, mirroring __hookPayloadBuffer) and settle the promise at construction — the single-scan continuation path previously dropped them and the await never settled Validated: core 1888 passed, full e2e 136/136 under WORKFLOW_VM=quickjs (nextjs-turbopack dev, world-local).
…-snapshots Carries the #3049 merge down the stack. Resolution notes: - inline loop's post-batch feed keeps BOTH sides: zero-event feeds raise pendingRequeueSignal before breaking (ack-without-requeue fix from #3049) and non-empty feeds advance eventsProcessedSinceSnapshot - quickjs-runtime.test.ts keeps both trailing suites (VM snapshot/restore + hook dispose-then-sleep replay) Validated: core 1899 passed; hook (26/26) and promiseRace e2e green under WORKFLOW_VM=quickjs with WORKFLOW_SNAPSHOT_THRESHOLD=5.
|
Re-ran against the latest head (6af2c2c) with the harness run-timeout raised 240s → 600s so QuickJS runs get a fair chance to complete, plus a variant with the newest stack layer (#3263's
Performance (postgres, from event-log forensics):
Takeaways: (1) the corruption repro is alive on this branch under node; (2) a QuickJS corruption comparison is not yet meaningful — completion rate is 0/144 across every QuickJS condition even at 600s; (3) the ~8× interpreter gap is the first-order blocker, and the new serde layer (as merged here) introduces a wake/dormancy bug worth catching before the restack. Happy to share the harness setup ( |
# Conflicts: # packages/core/src/runtime/quickjs-entrypoint.ts # packages/core/src/runtime/quickjs-runtime.ts # packages/world-postgres/src/drizzle/migrations/meta/_journal.json # packages/world-vercel/src/trace-propagation.test.ts
Port threshold snapshotting onto the post-#3263/#3342 engine: - Snapshot format v2: per-run snapshot metadata now carries the host-serde capture root's snapshot-portable token (serdeRootPtr) and the monotonic correlation-id ULID factory's last output (lastUlid). v1 snapshots (in-VM serde era) are treated as a miss. - Restore path rebuilds the host-side serde by re-adopting the capture root from the restored memory image (adoptSerdeRoot) — no guest code executes after user code has run, same mechanism as baseline snapshots. - ULID continuation: the factory's monotonic state lives host-side and does not survive into the memory image; restores continue the sequence via incrementUlidRandom (byte-equivalent to the ulid package's same-timestamp increment, asserted by test), so a restored invocation emits the exact ids full replay would — keeping concurrent invocations dedupable via per-(runId, correlationId) uniqueness. - __generateUlid joins the shared host-callback list (host functions are name-referenced from the WASM heap and must be re-registered on restore). - Preloaded event logs (no cursor) are ignored when restoring from a snapshot; the caller-attested complete preload fast path is kept for the no-snapshot paths.
Event Log Race Repro4 of 14 latest repro runs hit event-log regressions. Run History
Latest Scenario Breakdown
Latest Non-Completed Runs
|
…ent delete, gated diagnostics - @workflow/world: snapshots interface is now OPTIONAL on Storage (a World that can't store blobs simply omits it; consumers feature-detect and fall back to full replay) — keeps this a true minor for community worlds. New encodeSnapshotEnvelope/decodeSnapshotEnvelope helpers pack metadata + bytes into ONE self-describing blob; decode validates via the schema (passthrough for forward compat) and returns null for anything torn/corrupt/foreign — never fabricated metadata. - world-local: single envelope file per run replaces the .bin/.json pair — the torn-save window (crash or concurrent load between two renames pairing bytes with another suspension's cursor) is structurally gone. - world-postgres: data column stores the envelope; full metadata round-trips losslessly (new fields need no migration). Columns remain as denormalized observability copies. - world-vercel: envelope is the PUT/GET body, so the full metadata round-trips without any workflow-server change; loads never invent metadata from headers/wall time (undecodable body = clean miss). DELETE treats 404 as success (idempotent like local/postgres). WORLD_SNAPSHOT_DIAG logs gated behind the package's HTTP debug flag.
# Conflicts: # packages/world/src/index.ts
…ck restore, off-path compression, ceiling latch, optional snapshots - Blocking round-trip fix: full snapshot metadata now survives postgres and vercel via the snapshot envelope (see the world-side commit on quickjs-vm-snapshots) — the load gate no longer rejects every snapshot on production worlds. - world.snapshots is optional: the entrypoint feature-detects and forces the threshold to 0 when absent (pure full replay, always correct). - Engine-build pin: quickjs-wasi's version is baked into the generated assets module, stamped into snapshot metadata (engineVersion), and compared on load — the QJSS heap header is identical across builds, so a cross-build restore would execute as undefined behavior (live hazard mid-rollout). Mismatch = clean miss. - Deterministic clock high-water (clockMs) persists in metadata and primes the restored VM's clock — Date.now() no longer regresses to run-creation time until the first delta event. - Snapshot compression truly off the response path: the waitUntil IIFE yields past the current tick before touching bytes, and compress() gains preferAsync (zstd via the libuv threadpool with sync fallback; identical output bytes) so a multi-MB heap no longer blocks the loop. - restoredEventCount resets when a restore fails and the fallback refetches the full log — the ceiling no longer double-counts pre-snapshot events, and the next save no longer stamps an inflated eventCount. - Size-ceiling latch: a run whose heap once exceeded MAX_SNAPSHOT_PLAINTEXT_BYTES skips future captures outright (linear memory never shrinks), instead of re-paying two full heap copies per suspension to discard the result.
Combines the QuickJS-default dry-run with divergence-detection and write-fencing parity (#3453) so the event-log-race repro measures the same corruption classes on both engines. Conflict resolutions: - makeLiveSession takes both the snapshot-state accessor and the observed-events ledger; the snapshot-restore path runs the divergence sweep against its delta view (a throw there falls back to a fresh full replay, whose sweep is authoritative). - The precondition view tracker seeds its event count from the restored snapshot's metadata.eventCount (delta ids sort above the snapshot cursor, so watermark+count still describe the full log); snapshots without an eventCount fail open, and the restore-failure fallback rebuilds the view from the refetched full log.
Repro readout after merging quickjs-divergence-parity (#3453)The 2026-08-11 08:26 run is the first repro pass where the QuickJS engine carries the node engine's divergence detection + write fencing. The instrument now works, and the picture changes materially:
All four failures decrypt to the exact same signature as the node engine's hook-storm failures: i.e. the racing-writer / correlation-id-renumbering class — not a false positive of the new detection (a genuine one-off log would be reproduced by the recovery replays; these diverge identically on every recovery pass, which is the definition of an unsatisfiable log). What this means
The merge also closed the entrypoint's precondition-guard KNOWN GAP (all replay-context writes now fenced, including snapshot-restore-aware view seeding from |
Warning
Not intended to merge. This PR exists to exercise the full CI matrix — and the
event-log-race-reprostress workflow (via its label) — with QuickJS as the default engine, as a dress rehearsal for an eventual real default flip. It stays a draft; when the underlying stack (#3048 → #3049 → #3250 → #3251, plus #3263) is merged, a fresh flip PR will be cut against main with the learnings from this one.What it changes (for the dry run)
useQuickJSVmdefaults to the QuickJS engine when neither the run's stampedexecutionContext.workflowVmnorWORKFLOW_VMspecifies one;WORKFLOW_VM=nodeis the explicit node:vm opt-in.WORKFLOW_VMunset (exercising the default-selection path end to end); node legs opt in explicitly. Labels/artifacts unchanged viaMATRIX_VM.WORKFLOW_VM=node(see the comment at each pin).Findings so far (feeding the real flip PR)
event-limitred) was real and is fixed on QuickJS engine: inline step execution + WASM module caching #3049 (in-loop re-check + run_failed/MAX_EVENTS_EXCEEDED conversion).crypto, frozenprocess.env, loudIntl/locale guards); remaining differences are enumerated in the docs.start()— under a flipped default, unsetWORKFLOW_VMcurrently produces unstamped runs that would switch engines mid-run on rollback (review finding on this PR). The real flip PR must stamp the resolved engine.