Summary
drainDetachQueue's WebSocket recv-pause path cancels by raw fd with IORING_ASYNC_CANCEL_FD | IORING_ASYNC_CANCEL_ALL, which matches on req->file and cancels every in-flight op on the socket — not just the recv it intends. handleSend has no -ECANCELED case, so a cancelled SEND is treated as a fatal send error and the healthy connection is closed mid-write.
This is the same failure shape as #470 — a healthy connection destroyed by a cancel aimed at a different op — on a route the #470 generation fix cannot touch, because CANCEL_FD ignores user_data entirely.
Found by adversarial review of the #470 fix; survived 3/3 refutation votes.
Mechanism
engine/iouring/worker.go:3419-3423:
if cs.fixedFile {
prepCancelUserDataSkipSuccess(sqe, encodeUserDataGen(udRecv, cs.fd, cs.generation))
} else {
prepCancelFDSkipSuccess(sqe, cs.fd) // CANCEL_FD | CANCEL_ALL — generation-blind
}
engine/iouring/worker.go:2222 — the send completion has no cancellation branch:
if c.Res < 0 {
w.errCount.Add(1)
...
w.closeConn(fd)
}
Concrete sequence:
- A detached WebSocket conn has a broadcast frame in flight —
flushSend set cs.sending=true and the SEND is poll-armed because the peer's socket buffer is full.
- The
chanReader crosses its high-water mark and calls PauseRecv.
- The next
drainDetachQueue pass submits CANCEL_FD|CANCEL_ALL on cs.fd.
- The kernel cancels both the recv and the poll-armed send.
- The recv's
-ECANCELED is absorbed by the recvPaused guard (worker.go:1647). The send's is not — it lands in the c.Res < 0 branch and closes the connection mid-broadcast.
The failure is likeliest exactly when it matters most: a send is cancellable precisely when it is poll-armed on a full socket buffer, which is the same backpressure condition that triggers PauseRecv.
Note cs.fixedFile is forced false for transplant-adopted conns (transplant.go:92) regardless of worker capability, and fixed files are disabled at runtime on kernels that reject ACCEPT_DIRECT — so the fd-keyed branch is the common path, not the exotic one.
Why it was not caught by the #470 work
I instrumented this exact call site while chasing #470 and measured zero submissions (cancels=0 across 10 cells). That was correct but workload-specific: the observability refapp has no WebSocket route, so PauseRecv never fired. The trigger needs WS backpressure — ws-hub-broadcast territory, not h2c churn.
Suggested fix
Delete the cs.fixedFile branch and always use the generation-keyed form:
prepCancelUserDataSkipSuccess(sqe, encodeUserDataGen(udRecv, cs.fd, cs.generation))
The block's own comment (worker.go:3412-3414) already states the reason — user_data "is unambiguous in both modes" — and cancelConnOps (worker.go:2505) applies it unconditionally for exactly that reason. The fd-keyed branch is legacy and strictly over-broad: it cannot target the recv alone.
Deliberately not bundled into #481, which is scoped to the #470 root cause and must stay minimal for review.
Suggested verification
ws-hub-broadcast under backpressure (slow/stalled reader forcing a poll-armed send) with concurrent WS churn, on io_uring. Watch for connections closed with errIORingSend(-125).
Summary
drainDetachQueue's WebSocket recv-pause path cancels by raw fd withIORING_ASYNC_CANCEL_FD | IORING_ASYNC_CANCEL_ALL, which matches onreq->fileand cancels every in-flight op on the socket — not just the recv it intends.handleSendhas no-ECANCELEDcase, so a cancelled SEND is treated as a fatal send error and the healthy connection is closed mid-write.This is the same failure shape as #470 — a healthy connection destroyed by a cancel aimed at a different op — on a route the #470 generation fix cannot touch, because
CANCEL_FDignoresuser_dataentirely.Found by adversarial review of the #470 fix; survived 3/3 refutation votes.
Mechanism
engine/iouring/worker.go:3419-3423:engine/iouring/worker.go:2222— the send completion has no cancellation branch:Concrete sequence:
flushSendsetcs.sending=trueand the SEND is poll-armed because the peer's socket buffer is full.chanReadercrosses its high-water mark and callsPauseRecv.drainDetachQueuepass submitsCANCEL_FD|CANCEL_ALLoncs.fd.-ECANCELEDis absorbed by therecvPausedguard (worker.go:1647). The send's is not — it lands in thec.Res < 0branch and closes the connection mid-broadcast.The failure is likeliest exactly when it matters most: a send is cancellable precisely when it is poll-armed on a full socket buffer, which is the same backpressure condition that triggers
PauseRecv.Note
cs.fixedFileis forced false for transplant-adopted conns (transplant.go:92) regardless of worker capability, and fixed files are disabled at runtime on kernels that rejectACCEPT_DIRECT— so the fd-keyed branch is the common path, not the exotic one.Why it was not caught by the #470 work
I instrumented this exact call site while chasing #470 and measured zero submissions (
cancels=0across 10 cells). That was correct but workload-specific: theobservabilityrefapp has no WebSocket route, soPauseRecvnever fired. The trigger needs WS backpressure —ws-hub-broadcastterritory, not h2c churn.Suggested fix
Delete the
cs.fixedFilebranch and always use the generation-keyed form:The block's own comment (
worker.go:3412-3414) already states the reason — user_data "is unambiguous in both modes" — andcancelConnOps(worker.go:2505) applies it unconditionally for exactly that reason. The fd-keyed branch is legacy and strictly over-broad: it cannot target the recv alone.Deliberately not bundled into #481, which is scoped to the #470 root cause and must stay minimal for review.
Suggested verification
ws-hub-broadcastunder backpressure (slow/stalled reader forcing a poll-armed send) with concurrent WS churn, on io_uring. Watch for connections closed witherrIORingSend(-125).