fix(iouring): make conn generations process-unique (#470) - #481
Conversation
The generation stamped into a conn-bound SQE's user_data is the only thing
distinguishing one occupant of an fd from the next. acquireConnState
incremented a field on the POOLED connState:
cs := connStatePool.Get().(*connState)
cs.generation++
connStatePool is a sync.Pool, and GC drains it under connection churn, so
almost every acquire returns a freshly allocated connState whose generation
goes 0 -> 1. Measured on the validation workload: 973/973 connections and
628/628 close-path cancels carried gen=1. The generation provided ZERO
disambiguation between successive occupants of an fd -- the one thing it
exists to do.
cancelConnOps submits an ASYNC_CANCEL keyed on the recv's user_data
(udRecv, fd, gen). When the fd is recycled before the kernel runs that
cancel, the key matches the NEXT connection's recv byte-for-byte.
staleConnCQE sees the generations agree, accepts the CQE as that
connection's own (clearing recvArmed and decrementing kernelInflight), and
handleRecv treats the resulting -ECANCELED as a fatal read error and closes
a healthy connection that has never been read.
staleConnCQE documents this as a KNOWN RESIDUAL at "1/65536 per reuse".
That estimate assumes generations are spread across the 16-bit space. They
are not: they are pinned at 1, so the collision probability is ~100% on any
fd reuse that lands inside the cancel window.
Proven by connection identity, not inference:
fd=130 live_cid=8255 killed by a cancel from cid=6241, 80us earlier
fd=138 live_cid=23059 killed by a cancel from cid=21043, 65us earlier
both with gen 1 == 1.
On the wire (tcpdump, loopback): the server ACKs the client's full 126-byte
h2c request, sends ZERO response bytes, then FIN followed by RST 273us
after SYN. The RST is the kernel's response to close() with a non-empty
receive queue -- the request was never read.
Fix: draw the generation from a process-monotonic counter instead of the
pooled object's own field, skipping the reserved gen==0 (encodeUserDataGen
collapses gen=0 onto the plain encodeUserData encoding). A collision now
requires 65536 intervening accepts on the same fd, which the
microsecond-wide cancel window cannot span.
No behavioural change beyond the generation's source: one atomic increment
per accepted connection, on a path that already performs accept,
getpeername and setsockopt syscalls. No new branches in the hot path.
engine/epoll has its own acquireConnState and is untouched.
Validation on the cluster (real cmd/validator, C=120, amd64):
with fix: 0 hangs / 24 cells / 232,128 read conns
fix reverted: 4 hangs / 6 cells / 58,032 read conns
P(0 | pre-fix baseline) ~ 9e-14.
Explains the full history: 27/27 h2c_hang events across 8 nightlies were
io_uring, 0 epoll, 0 std (p ~ 5e-13) -- ASYNC_CANCEL keyed on user_data has
no epoll/std analogue.
Records the cost of the change in conn.go: 15.45 -> 16.22 ns/op on acquireConnState alone (+0.77 ns, 0 allocs), which is 0.0025% of a ~31us connection lifecycle and unmeasurable end to end (1,437,649 vs 1,438,729 conns per 45s across alternating trials).
Follow-up to the process-monotonic generation, from an adversarial review of that fix. Two of its claims were wrong and the comment I wrote asserted them: 1. The wrap is PROCESS-WIDE, not per-fd. connGenSeq is drawn once per accepted fd across every worker, so a wrap costs 65536 accepts ANYWHERE. At this cluster's measured end-to-end rate (1,437,649 conns / 45s = ~31.9k accepts/s) a 16-bit generation wrapped every 2.05 SECONDS. 2. The binding window is NOT the microsecond cancel latency. It is the paths where no cancel is submitted at all: cancelConnOps skips the ASYNC_CANCEL when the SQ ring is full (worker.go:2504 has no else branch), leaving the op to the 5s pendingRelease backstop; and an armed udHeaderTimer lives for ReadHeaderTimeout, 10s by default. Those windows spanned 2.4 and 4.9 wraps of a 16-bit generation. So the 16-bit fix reduced the #470 misroute by 3-5 orders of magnitude (consistent with 0 hangs in 348,192 connections) but did not close it. The residual per stranded op was ~1.3e-3 at the validator's C=120 shape. The generation field was 16 bits only because the fd field was given 40 -- and the layout comment already noted fds and fixed-file indices are "both < 65536 in practice". onAcceptedFD in fact REJECTS any fd >= fixedFileTableSize (65536) outright, so 24 bits of fd is 256x the engine's own hard bound. Re-splits user_data as op(56-63) | gen(24-55, 32 bits) | fd(0-23, 24 bits). A collision now needs 2^32 intervening accepts -- ~37 hours at 31.9k/s -- which no engine-side window can span. Same shift-or on the hot path: no added logic, no measurable cost. Also corrects the over-claiming comment in acquireConnState, and strengthens the encoding tests: gens now span the full 32-bit range including the old 16-bit ceiling, and TestGenerationDiffersAcrossReuse asserts uniqueness UNCONDITIONALLY. Its previous version could not -- it noted that "a fresh pool object starts at generation 1, colliding with a once-used object's 1" and worked around it. That comment was describing celeris#470 as a test flake.
Verification completeCluster validationReal
502,944 connections with the fix, zero events. Reverting reproduces the bug immediately. Full-breadth nightlyprobatorium run 33771655960, on a branch pinning all 8 refapps to this commit — 48 cells (8 refapps x 3 engines x 2 arches), 0 hangs, 0 eof, 0 timeout. Honest note on power: the nightly runs at the default Correctness
Performance
Inside the baseline's own 0.28% trial spread. At ~31 us per connection the atomic is 0.0025% of the path. The 32-bit widening adds nothing further — same shift-or. Second commit: 32-bit generationAn adversarial review of the first commit found my safety claim was wrong in both clauses, and I had written it into the code:
So the 16-bit fix cut the misroute by 3-5 orders of magnitude but left a residual (~1.3e-3 per stranded op). The generation was 16 bits only because the fd field had 40 — and Related, deliberately NOT in this PR#482 — the WS recv-pause path cancels by raw fd ( Note for the mergerprobatorium branch |
Fixes #470.
Root cause
The generation stamped into a conn-bound SQE's
user_datais the only thing distinguishing one occupant of an fd from the next.acquireConnStateincremented a field on the pooledconnState:connStatePoolis async.Pool, and GC drains it under connection churn — so almost every acquire returns a freshly allocatedconnStatewhose generation goes0 → 1. Measured on the validation workload:The generation provided zero disambiguation between successive occupants of an fd, which is the one thing it exists to do.
cancelConnOpssubmits anASYNC_CANCELkeyed on the recv'suser_data—(udRecv, fd, gen). When the fd is recycled before the kernel runs that cancel, the key matches the next connection's recv byte-for-byte.staleConnCQEsees the generations agree, accepts the CQE as that connection's own (clearingrecvArmed, decrementingkernelInflight), andhandleRecvtreats the resulting-ECANCELEDas a fatal read error and closes a healthy connection that has never been read.staleConnCQEdocuments this as a KNOWN RESIDUAL at "1/65536 per reuse". That estimate assumes generations are spread across the 16-bit space. They are pinned at 1, so the real probability is ~100% for any fd reuse landing inside the cancel window.Evidence
By connection identity (unique per-conn ids, not inference):
On the wire (tcpdump, loopback):
ack=127— the server ACKed the client's full 126-byte h2c request.seq=1, len=0— it sent nothing. RST after FIN is the kernel's response toclose()with a non-empty receive queue: the request was never read. 273 µs, mid-run.The fix
Draw the generation from a process-monotonic counter instead of the pooled object's own field, skipping the reserved
gen==0(encodeUserDataGencollapsesgen=0onto the plainencodeUserDataencoding). A collision now requires 65,536 intervening accepts on the same fd, which the microsecond-wide cancel window cannot span.One production file, four functional lines. No new branches in the hot path, no behavioural change beyond the generation's source.
engine/epollhas its ownacquireConnStateand is untouched.Verification
Run with the real
cmd/validatoragainst a single refapp,VALIDATE_CONCURRENCY=120, per-cell parameters bit-identical to a production nightly cell.engine/iouringsuite on the SUTgeneration 1 reused (acquire #0 and #1))P(0 hangs | pre-fix baseline of 1.25/cell over 24 cells) ~ 9e-14.Reverting the fix brings the bug straight back — causality, not correlation.Performance
acquireConnStatemicrobenchThe end-to-end difference sits inside the baseline's own 0.28% trial spread. At ~31 µs per connection the atomic is 0.0025% of the path.
Why this explains the whole history
27/27
h2c_hangevents across 8 nightlies were io_uring — 0 epoll, 0 std (p ≈ 5e-13).ASYNC_CANCELkeyed onuser_datahas no epoll/std analogue. The event is load-correlated (4× walker fan-out → 5× events) because more churn means more fd reuse inside the cancel window.Notes for review
uint16and wraps. A collision needs 65,536 intervening accepts on the same fd — far outside the cancel window — but that is a bound, not a proof. A wider field would needuser_datalayout changes; out of scope here.handleRecvstill treats-ECANCELEDas fatal. With unique generations only a connection's own cancel can match, so that is now correct. I deliberately did not add defensive handling — it would be unneeded logic on the hot path.