QuickJS engine: threshold-based VM-memory snapshotting (WORKFLOW_SNAPSHOT_THRESHOLD) - #3053
Merged
Merged
Conversation
🦋 Changeset detectedLatest commit: f1b1e10 The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
Contributor
Contributor
🧪 E2E Test Results✅ All tests passed E2E Test SummarySummary
Details by Category✅ ▲ Vercel Production
✅ 💻 Local Development
✅ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ 📋 Other
✅ vercel-multi-region
|
TooTallNate
force-pushed
the
quickjs-vm-snapshots
branch
from
July 31, 2026 01:07
1e0b6e6 to
f43ee63
Compare
TooTallNate
force-pushed
the
quickjs-vm-threshold-snapshots
branch
from
July 31, 2026 01:09
64f40c3 to
f1b1e10
Compare
TooTallNate
force-pushed
the
quickjs-vm-snapshots
branch
from
July 31, 2026 02:15
f43ee63 to
f64932e
Compare
Thegreatsura
pushed a commit
to Thegreatsura/workflow
that referenced
this pull request
Oct 2, 2026
…SHOT_THRESHOLD) (vercel#3251) > [!NOTE] > Supersedes vercel#3053. Stacked on vercel#3250 (`quickjs-vm-snapshots`). Review only the commits after vercel#3250's. ## Summary Experimental **threshold-based VM-memory snapshotting** for the QuickJS engine. Snapshots are taken only once `WORKFLOW_SNAPSHOT_THRESHOLD` events have been processed since the last one, not at every suspension: - **Short-lived runs never snapshot.** They keep pure replay with no snapshot round-trips. - **Long/forever runs stop scaling their resume cost with event-log length.** A resumption restores the VM heap and replays only the events recorded after the snapshot's cursor. ## How it works - **Policy:** the `WORKFLOW_SNAPSHOT_THRESHOLD` env var (default `0`, disabled) or per-run `executionContext.snapshotThreshold`, stamped at `start()` like `WORKFLOW_VM`. An invalid handler-side value disables snapshotting with a warning. - **Save** (suspension exit, threshold met): capture the VM memory. Then, after the response goes out (`waitUntil`): 1. Frame the heap with its restore-relevant metadata, bound to the run id. 2. Compress (zstd on the threadpool, gzip fallback) and encrypt with the run's key. 3. `world.experimental_snapshots.save`. Captures over 32 MB plaintext are skipped, and the run is latched so later suspensions skip the capture. - **Restore** (later invocation): 1. `load`, check format/engine version and bounds on the metadata. 2. Decrypt, and decompress with a size cap. 3. Verify the sealed metadata matches the envelope. The saved position is the read cursor plus the exact number of events it covers, so a restore adds only the events listed after it. 4. Restore over the cached WASM module, re-register every host callback from one shared list, and reinstall `process.env` from the current invocation. 5. Replay the delta. The load is skipped when no snapshot can exist yet (log below the threshold). - **Encryption:** runs without an encryption key are only snapshotted with `WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED=1`. A run with a key only accepts snapshots encrypted with it. - **Delete:** after the run's terminal event, off the response path, whenever a snapshot may exist. A save that lands after the run finished deletes itself. - **Fallback is always full replay.** A missing, rejected, or unrestorable snapshot logs a warning and boots fresh against the full log. The log remains the source of truth. ## Determinism model (restore + partial replay) A resumption may restore a snapshot **older** than the log head and must re-derive everything in between deterministically: - **Seeded PRNG:** it stays on the run's base seed. The snapshot records how many draws the heap consumed (`rngDraws`), and restore fast-forwards that many. Ids are therefore position-based: identical across snapshot generations, identical to a no-snapshot full replay, and identical across concurrent resumes from different snapshots, so the world's dedup still collapses them. - **ULID factory and clock:** the correlation-id ULID factory state (`lastUlid`) and the deterministic clock's high-water mark (`clockMs`) are persisted and continued. - **Re-fed events:** feeding already-consumed events is harmless. Settled resolvers are gone, re-scanned terminals for consumed ids are dropped, and hook deliveries are deduped by eventId. ## Validation - Unit tests with the real VM: - restore/resume, and id parity with full replay; - partial replay across multiple suspensions; - host-callback re-registration; - current `process.env` after restore; - no step results retained in the heap. - Differential fuzz (`quickjs-snapshot-fuzz.test.ts`): random parallel, sequential and raced step schedules driven live, through random restores (including older snapshots, lagging re-feeds and split bursts), and as one full replay, all of which must agree. A short run is in the unit suite; a nightly workflow runs more seeds and steps. - Real-VM entrypoint tests against an in-memory World (`quickjs-snapshot-generations.test.ts`): - `MaxEventsExceeded` fires exactly at the limit across many save/restore generations; - each saved cursor covers exactly its saved event count; - thresholds 1/3/4/1000 write the same log and result as no snapshots and leave nothing behind; - an older save landing after a newer one converges. - Entrypoint tests with snapshot storage mocked: - load gating; - delete on completion, on restore failure, and past the threshold; - rejection of plaintext, mismatched-metadata, other-run, out-of-bounds, and oversized snapshots. - CI: a `quickjs-snapshot` matrix leg (nextjs-turbopack, threshold=1) across the local dev/prod/postgres e2e jobs. It fails if the server log shows no snapshot restores. ## Notes - Snapshot bytes are tied to the `quickjs-wasi` build that produced them. A mismatch is a clean miss. On Vercel, runs stay on their deployment, so deploys don't invalidate in-flight snapshots. - Docs: `WORKFLOW_SNAPSHOT_THRESHOLD` and `WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED` in v5 Runtime Tuning, including what a snapshot contains and how it is protected. --------- Signed-off-by: Peter Wielander <mittgfu@gmail.com> Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> Co-authored-by: Nathan Rajlich <71256+TooTallNate@users.noreply.github.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
PR 4 of the QuickJS VM roadmap: threshold-based VM-memory snapshotting — the middle ground that motivated reviving this effort (see #1298 / #1300 discussion). Instead of snapshotting at every suspension (the original branch's model, which cost ~25% on e2e wall clock), snapshots are taken only once
WORKFLOW_SNAPSHOT_THRESHOLDevents have been processed since the last one:How it works
WORKFLOW_SNAPSHOT_THRESHOLDenv var (default0= disabled) or per-runexecutionContext.snapshotThreshold, stamped atstart()for run affinity likeWORKFLOW_VM.session.snapshot()) → compress (zstd/gzip via the shared serialization pipeline; QuickJS heaps compress ~4×, measured 16.5 MB → 3.9 MB) → encrypt with the run's key when configured →world.snapshots.savewith the events cursor at the VM's feed frontier.world.snapshots.load→ decrypt → decompress →QuickJS.restoreover the cached WASM module, re-register host callbacks, fetch events from the snapshot's cursor and feed only the delta. Runs seamlessly through PR 2's inline continuation loop.Determinism model (restore + partial replay)
The threshold model's new mechanism vs. the original branch: a resumption may restore a snapshot older than the log head (suspensions since the snapshot weren't persisted) and must deterministically re-derive everything in between:
eventsCursor: the heap already consumed pre-snapshot draws, so re-seeding from the base would replay the first-N draws and collide with recorded correlationIds. The cursor is identical for every resume from the same snapshot (concurrent resumes still collide ids for the world's dedup) and advances only when a newer snapshot is taken.Validation
WORKFLOW_SNAPSHOT_THRESHOLD=1(maximum churn: snapshot on every qualifying suspension), wall clock within ~10% of the node baselinerestored: trueresumptions, save/restore/delete lifecycle, and threshold gating (threshold=100 short run ⇒ zero snapshots, pure replay)quickjs-snapshotmatrix leg (nextjs-turbopack, threshold=1) across local dev/prod/postgres e2e jobsNotes / follow-ups
quickjs-wasibuild that produced them. Per the deployment contract (runs continue on the version they started on — free on Vercel), this is a non-issue in production; environments without skew protection are covered by the restore-failure fallback to full replay.WORKFLOW_SNAPSHOT_THRESHOLDsection added to v5 Runtime Tuning.