Repository navigation
[world-vercel] Raise H2 receive windows on the events agent - #3212
Conversation
🦋 Changeset detectedLatest commit: 4627aa3 The changes in this PR will be included in the next version bump. This PR includes changesets to release 17 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests📦 Local Production (1 failed)nextjs-webpack-stable (1 failed):
E2E Test SummarySummary
Details by Category✅ ▲ Vercel Production
✅ 💻 Local Development
❌ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ 📋 Other
✅ vercel-multi-region
|
📊 Workflow Benchmarkscommit Backend:
📜 Previous results (1)b2d0939Thu, 30 Jul 2026 18:25:10 GMT · run logs
ℹ️ Metric definitions & methodologyBest/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120 All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor ( Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the |
undici defaults to a 256 KiB stream window and a 512 KiB connection window, both far below the bandwidth-delay product of the function-to-edge path (~9 ms RTT measured from iad1). Event-log reads on replay exceed them, so the origin stalls for a round trip per window's worth of body. Both windows have to be raised together: a single read is capped by the stream window, concurrent reads share the connection window, and raising either alone leaves the other binding. Measured against a loopback H2 origin mirroring the edge's SETTINGS at 10 ms RTT, 4 MiB/16 MiB cuts read wall time 78-89% across 1, 4 and 32 concurrent reads, against 1.8-2.6% noise floors. This also removes an interaction with multiplexing: concentrating streams onto one connection makes them share its receive window, so before this change pipelining=1 measured ~70% faster than pipelining=100 on downloads purely because 8 connections gave 8 separate windows. The write path is unaffected — receive windows do not govern uploads. maxConcurrentStreams is left alone; it is inert, since undici only uses it to seed peerMaxConcurrentStreams at connect and the server's SETTINGS then overwrite it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
b2d0939 to
f28e947
Compare
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
karthikscale3
left a comment
There was a problem hiding this comment.
Reviewed the HTTP/2 receive-window tuning and its tests. The options and flow-control behavior match the pinned Undici 7.28 implementation, the change is correctly scoped to the events agent, and I found no blocking correctness or regression issues.
|
No backport to This is a performance optimization (tuning HTTP/2 receive flow-control windows to reduce read latency), not a fix for a correctness, crash, or resource-leak defect on To override, re-run the Backport to stable workflow manually via |
Follow-up to #3190. That PR made the events agent actually multiplex; this one tunes the flow-control windows it runs with.
Problem
undici defaults to a 256 KiB stream window and a 512 KiB connection window (
lib/dispatcher/client.js:276-277). Both are far below the bandwidth-delay product of the function→edge path — a Function iniad1measures ~9 ms RTT to the edge (min 6.2, p50 9.4–9.8, against edge-terminated routes). Once a response exceeds the window the origin stops and waits for aWINDOW_UPDATE, costing a round trip per window's worth of body. Event-log reads on replay are exactly that shape: pages are read sequentially, page size is server-driven with no byte cap, and step payloads are serialized into events, so a page can be several MiB.There's also an interaction with #3190 worth calling out: concentrating streams onto one connection makes them share that connection's single receive window. Before this change,
pipelining: 1measured ~70% faster thanpipelining: 100on downloads, purely because spreading over 8 connections gave 8 separate windows. Multiplexing was paying for itself on writes and giving some back on reads.Change
EVENTS_AGENT_OPTIONSgainsinitialWindowSize: 4 MiBandconnectionWindowSize: 16 MiB. Nothing else changes — the default (H1) and stream-write agents are untouched.Both windows have to move together. Whichever is left at its default becomes the binding constraint on its own (rtt=10ms):
initialWindowSizealoneconnectionWindowSizealoneA single read is capped by the stream window, so raising only the connection window does nothing. Concurrent reads share the connection window, so raising only the stream window relocates the stall and measures slightly worse than leaving both alone.
Results
Shipped config vs undici defaults, duplicate pairs agreeing within 0.5%:
Sizing is the knee, not a round number: 4 MiB of stream window captures the single-read win in full (216 → 17 ms; 2 MiB only reaches 29 ms), and 16 MiB of connection window is needed for the concurrent case (8 MiB is no better than 1/8 at conc=32, 16 MiB is 25% faster, 32 MiB adds nothing).
maxConcurrentStreamsis deliberately not set. It's inert: undici only uses it to seedpeerMaxConcurrentStreamsat connect, after which the server's SETTINGS overwrite it (the edge advertises 160). Measured at ±1.6% inside a 6.1% noise floor across every shape.pipeliningremains the only knob that gates in-flight streams.Costs
Methodology
Numbers come from a local harness, kept out of this PR: a loopback H2 origin mirroring the edge's advertised SETTINGS (
MAX_CONCURRENT_STREAMS160,INITIAL_WINDOW_SIZE1 MiB,MAX_FRAME_SIZE1 MiB) behind a latency-injecting TCP relay. The relay is not optional — loopback RTT is ~0.05 ms, at which every flow-control window is thousands of times larger than the BDP, so these knobs are unmeasurable by construction without injected RTT.Two false positives shaped the design, and both are easy to repeat:
connectionWindowSize=16MiBresult of −31.9% that did not replicate at all. Separately, a 3-baseline write-path run showed a "reproducible" +3.2% regression against a 2.0% floor that vanished when a fourth control put the floor at 4.8%.Every reported number therefore comes from: interleaved duplicate controls (the spread is the noise floor), round-robin execution with seeded reshuffling each round (an earlier sequential version produced results that improved monotonically down the list — the first config was absorbing warmup), pre-warmed dispatchers, and median-of-rounds.
Draft because the numbers above are from the local rig; leaving it open for CI and review before marking ready. Happy to share the harness separately if it's useful for re-deriving these constants.
🤖 Generated with Claude Code