[core] Serialize only the viewed bytes of a Float16Array - #4594
Conversation
🦋 Changeset detectedLatest commit: b61f19b The changes in this PR will be included in the next version bump. This PR includes changesets to release 17 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed 🛠 Infra Events (absorbed by the harness)Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.
E2E Test SummarySummary
Details by Category✅ ▲ Vercel Production
✅ 💻 Local Development
✅ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ 🌐 Cross-language Conformance
✅ dynamic-runs
✅ vercel-http-transport
✅ vercel-multi-region
✅ vercel-ws-transport
|
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 146558ms → this run 209613ms (Δ +63055ms, +43%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The canonical QuickJS reference codec still emits the unsafe legacy representation, and supported-type documentation remains outdated.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Prevents Float16Array subviews from serializing unrelated backing-buffer bytes and adds equivalent QuickJS support.
Changes:
- Copies only viewed bytes before devalue serialization.
- Adds QuickJS
Float16Arraydetection and revival. - Adds cross-engine and pooled-buffer tests.
| File | Description |
|---|---|
serialization/hardened.ts |
Restricts serialized views to their byte range. |
serialization.test.ts |
Tests pooled, whole, and empty views. |
runtime/quickjs-serde.ts |
Adds QuickJS Float16Array support. |
runtime/quickjs-serde.test.ts |
Tests parity and legacy revival. |
.changeset/unclaimed-view-pool-leak.md |
Records the patch release. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| // The node:vm engine's encoding is the contract: `Float16Array` has no | ||
| // workflow reducer on either engine, so both must emit devalue's built-in | ||
| // `[tag, buffer]` form, and a subview must carry only its viewed bytes. | ||
| const nodeEngine = (value: unknown) => text(nodeEngineSerialize(value)); |
There was a problem hiding this comment.
Agreed, fixed in b61f19b. codec-devalue-vm.ts now passes a viewInfo operation that copies a subview's viewed bytes into a fresh buffer, the same as hardened.ts and quickjs-serde.ts. The Float16Array parity tests now assert that QuickJS output is byte-identical to both referenceSerialize and the node:vm engine. Without the reference-codec change, the subview and empty-subview cases fail.
| "@workflow/core": patch | ||
| --- | ||
|
|
||
| Serialize only the viewed bytes of a `Float16Array` (no longer the whole Node `Buffer` pool behind it), and support `Float16Array` in the QuickJS engine with the same wire format as the node:vm engine. |
There was a problem hiding this comment.
Agreed for v5, fixed in b61f19b: Float16Array is now listed next to Float32Array/Float64Array in docs/content/docs/v5/foundations/serialization.mdx, and the PR description has a Docs Preview section.
I left v4 alone on purpose. Nothing in this PR applies to 4.x: stable has no QuickJS engine and no hardened viewInfo (it calls devalue's stringify with the default operations). Whether to list Float16Array in the v4 docs belongs with a fix on stable, not here.
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 42 total
Full trace: |
About these numbersSizes are gzip; parentheses show the change against
|
|
No backport to The fix lands almost entirely in code that does not exist on To override, re-run the Backport to stable workflow manually via |
## Summary `packages/core/src/runtime/vm-serde-bundle.generated.ts` is dead code that was committed by accident. - #3048 generated it at build time with `scripts/build-vm-serde-bundle.js` and gitignored it. It was the serializer bundle evaluated *inside* the QuickJS VM. - #3263 moved QuickJS serialization to the host (`runtime/quickjs-serde.ts`). That PR deleted the build script and the gitignore entry, but the leftover generated file got committed in the same change. Since then: - nothing imports it, and no build step regenerates it (core's `turbo.json` outputs list only `quickjs-assets.generated.ts`); - `quickjs-serde.test.ts` refers to it as "the retired in-VM serde bundle"; - it is stale: it predates #4443, so its reducers have no `DataViewBytes`. It never runs, but `tsc` compiles it into `@workflow/core`'s `dist`. It is also easy to mistake for live serializer code during an audit (it came up while auditing #4594). ## Changes - Delete the file. - Drop its path from the e2e build-artifact list in `tests.yml`, which is meant to stay aligned with the `turbo.json` build outputs. ## Testing - `tsc --noEmit` reports nothing referencing it; `src/runtime` tests pass (1256).


Summary
Node allocates small
Buffers (Buffer.allocUnsafe, smallBuffer.from) as windows onto a shared 8 KiB pool. A serializer that encodes a view's whole backingArrayBuffertherefore writes unrelated, possibly sensitive pool memory into the run's event log. devalue is shipping a fix for this in its defaultstringifyoperations; this PR is our audit of the same class.Audit
Every typed array except one was already safe: the workflow reducers claim
Buffer/Uint8Array, the other typed arrays, and (since #4443)DataView, and encode only the viewed range. Probed through the real codec with a pooled Buffer neighbouring a secret:Buffer, typed arrays,DataViewFloat16Arrayover a pooled Bufferbuf.bufferpassed explicitlyFloat16Arrayhas no reducer, so it reaches devalue's built-in typed-array branch, which calls our hardenedviewInfooperation. That returned the whole backing buffer plus offset/length. Upgrading devalue would not fix this, because we overrideviewInfo.Changes
serialization/hardened.ts): the hardenedviewInfocopies a subview's viewed bytes into a fresh buffer (through captured intrinsics, so guest patches can't influence it) and reports it as a whole-buffer view. devalue then emits the compact["Float16Array", buffer]form, which every existing reader already parses. Whole-buffer views are unchanged.serialization/codec-devalue-vm.ts, the value-space model the QuickJS wire fixtures are checked against): same subview copy, so it agrees with both engines.runtime/quickjs-serde.ts):Float16Arraypreviously classified as a plain object and threwCannot stringify arbitrary non-POJOs(no leak: the VM has noBufferpool). It's now sampled and captured, so it takes the same devalue built-in branch, and the QuickJSviewInfoapplies the same subview copy, so both engines emit byte-identical payloads. The constructor capture is optional so a capture root adopted from an older snapshot still works.Why not add a
Float16Arrayreducer (or drop the existing ones)?devalue skips its built-in branch for any tag that has a custom reviver, and hands that reviver only the hydrated referent. A
Float16Arrayreviver would therefore strip the bounds off["Float16Array", buf, off, len]payloads already in event logs (the same reason #4443 used theDataViewBytestag). Conversely, the existing typed-array reducers are the wire format for every stored["Uint8Array","<base64>"], so they have to stay.Tests
serialization.test.ts: pooledFloat16Arraystep return carries only its 4 viewed bytes and round-trips; whole-buffer and empty subview round-trip.quickjs-serde.test.ts: whole-buffer, subview and empty-subviewFloat16Arrayare byte-identical to both the reference codec and the node:vm engine and revive in the guest; a payload carrying subview bounds revives with them.src/runtime+src/serialization*pass locally (1888 tests).Docs Preview
v4 docs are unchanged: this PR's changes are 5.x-only (see review thread).