Found while working the review of #696 (celeris#691). Pre-existing on main (9f4d89b), and not changed by #696: the same probe fails 3/3 on both.
The defect
failDriverConn (a driver SEND fails while the conn's RECV is still armed) marks the close pending and queues an ASYNC_CANCEL whose user_data is udDriverClose|fd. Driver user_data carries no generation (encodeUserData, not encodeUserDataGen as HTTP conns use since #470), and handleDriverClose looks the conn up by that number.
When the RECV's own CQE is processed before the cancel's CQE, the RECV finalizes the conn: it leaves driverConns and onClose fires. The cancel's CQE is processed an iteration later. If the caller has closed the fd and registered a new conn that got the same number in between, the stale CQE is routed to the new conn:
- with an op in flight, the new conn is marked
closePending, and its first CQE finalizes it;
- with none, it is finalized at once.
Either way the new conn's onClose(nil) fires, and the data of its first RECV is dropped (the closePending check runs before onRecv).
Measured
A probe forces the order: the worker is parked (a driver onRecv blocks) so the next batch holds Q's RECV CQE and then A's; while parked in Q, failDriverConn(A) runs (A's RECV still counted in flight); A's RECV CQE then finalizes A, and A's onClose parks while the test closes A's fd, gets the same number for a new socket N, and registers N.
| head |
runs |
N's first byte reached onRecv |
N's onClose fired |
9f4d89b (main) |
3 |
0/3 |
3/3, with a nil error |
#696 at 5975527 |
3 |
0/3 |
3/3, with a nil error |
In every run the batch order was Q then A, and A had 1 op in flight at failDriverConn. Laptop Docker, golang:1.27, -race, memlock 8 MiB, one io_uring worker (kernel 7.0.12-linuxkit, DEFER_TASKRUN).
Evidence (maintainer's evidence tree): evidence/celeris-691/lane-20260926/round2/, probe patches/zz_probe691r2_staleclose_test.go, script tools/q9-staleclose.sh, logs logs/q9-01-r2sbase-m8.log and logs/q9-02-r2sfix-m8.log.
Who reaches it
A driver that registers a new conn from its onClose handling within about one worker iteration, and gets the same descriptor number, which is the norm once the old fd is closed (lowest free number).
Fix direction
Either one closes it:
- Count an issued driver cancel as in flight (
dc.cancels++ when it is prepared, -- in handleDriverClose), and finalize only when both the ops and the cancels have settled. A close CQE then never outlives its conn, and nothing else changes.
- Add a generation to driver
user_data, as HTTP conns have (encodeUserDataGen), and drop a close CQE whose generation does not match.
(1) is the smaller change. The probe above is the failing-first test for either.
Found while working the review of #696 (celeris#691). Pre-existing on
main(9f4d89b), and not changed by #696: the same probe fails 3/3 on both.The defect
failDriverConn(a driver SEND fails while the conn's RECV is still armed) marks the close pending and queues anASYNC_CANCELwhoseuser_dataisudDriverClose|fd. Driveruser_datacarries no generation (encodeUserData, notencodeUserDataGenas HTTP conns use since #470), andhandleDriverCloselooks the conn up by that number.When the RECV's own CQE is processed before the cancel's CQE, the RECV finalizes the conn: it leaves
driverConnsandonClosefires. The cancel's CQE is processed an iteration later. If the caller has closed the fd and registered a new conn that got the same number in between, the stale CQE is routed to the new conn:closePending, and its first CQE finalizes it;Either way the new conn's
onClose(nil)fires, and the data of its first RECV is dropped (theclosePendingcheck runs beforeonRecv).Measured
A probe forces the order: the worker is parked (a driver
onRecvblocks) so the next batch holds Q's RECV CQE and then A's; while parked in Q,failDriverConn(A)runs (A's RECV still counted in flight); A's RECV CQE then finalizes A, and A'sonCloseparks while the test closes A's fd, gets the same number for a new socket N, and registers N.9f4d89b(main)5975527In every run the batch order was Q then A, and A had 1 op in flight at
failDriverConn. Laptop Docker,golang:1.27,-race, memlock 8 MiB, one io_uring worker (kernel 7.0.12-linuxkit, DEFER_TASKRUN).Evidence (maintainer's evidence tree):
evidence/celeris-691/lane-20260926/round2/, probepatches/zz_probe691r2_staleclose_test.go, scripttools/q9-staleclose.sh, logslogs/q9-01-r2sbase-m8.logandlogs/q9-02-r2sfix-m8.log.Who reaches it
A driver that registers a new conn from its
onClosehandling within about one worker iteration, and gets the same descriptor number, which is the norm once the old fd is closed (lowest free number).pubsub.go:103,async.NewBackoff(50ms, 5s)) before it redials, so it hits this only when the worker is still inside the same batch after 50 ms, for example behind a slow inline handler (the io_uring: UnregisterConn then Close leaks the driver socket — the fd-keyed ASYNC_CANCEL misses once the caller has closed the fd, so onClose never fires and the peer never sees EOF #691 busy probe held workers for 400 ms).srv.EventLoopProvider().WorkerLoop(n)API that reconnects fromonClose.Fix direction
Either one closes it:
dc.cancels++when it is prepared,--inhandleDriverClose), and finalize only when both the ops and the cancels have settled. A close CQE then never outlives its conn, and nothing else changes.user_data, as HTTP conns have (encodeUserDataGen), and drop a close CQE whose generation does not match.(1) is the smaller change. The probe above is the failing-first test for either.