Found by reading 468ce53 during the celeris#639 fd audit. The leak is visible in the code: every step is cited below. Its rate under real load has not been measured.
The leak
Worker.run in engine/iouring/worker.go:
- Creates the worker's
SO_REUSEPORT listen socket and stores it: createListenSocket (815), w.listenFD = listenFD (820).
- If ring setup fails, it sends
ready <- fmt.Errorf("worker %d ring setup: …") (827) and returns without closing listenFD.
- If the first submit fails, it sends
ready <- fmt.Errorf("worker %d initial submit: …") (913) and returns without closing listenFD or the ring.
The worker's run has already returned, so its shutdown never runs. Engine.Listen (engine.go:261-265) cancels the other workers, joins them, and returns the error. Nothing ever closes the failed worker's listen socket.
Why this is worse than a leaked descriptor
The socket is still listening, in the port's SO_REUSEPORT group. The kernel keeps hashing a share of new connections to it. Nobody accepts from it, so those clients complete the handshake into a backlog that is never drained, and they hang.
The adaptive engine is where that group is shared.
- A lazy standby build (
adaptive/engine.go:685 for io_uring) starts built.Listen on the same port as the active epoll engine, then waits 5 s for Addr().
- When the io_uring Listen fails as above, the wait times out.
performSwitch logs aborting switch: lazy standby build failed; staying on current active and returns (686-689).
- Nothing records a switch attempt or starts a cooldown on that path, and
e.secondary stays nil. The eval ticker is 1 s (controller.go:95). Unless ctrl.evaluate backs off by itself (not checked here), the next tick builds and fails again.
- Each failed attempt leaves one more dead listener on the serving port. Each attempt also holds
e.mu for the full 5 s bind wait.
The engine is left serving on epoll with a growing fraction of its incoming connections stranded. No error comes back to the caller; only a Warn is logged.
What would make a worker's ring setup fail after the tier probe passed
The tier probe (engine.go ~220-238) creates and closes a single test ring. Per-worker rings are created later, one per worker.
The plausible causes are:
- memlock accounting across N rings;
- rings leaked by earlier failed attempts (the submit-failure path above leaks one);
SQPOLL limits.
The CI shape (RLIMIT_MEMLOCK 8 MiB, see #533's history) is where ring creation is known to be tight. Not yet measured.
Fix
- Close
listenFD (and, on the submit path, the ring) on every early return after it is created. A defer that runs unless the worker reached ready <- nil covers all of them.
- Add a test that injects a ring setup failure for one worker. It asserts that no listening socket for the port survives
Listen's error: count listeners on the port via /proc/net/tcp, or try bind without SO_REUSEPORT. The test must fail on main.
- Separately, decide whether an aborted lazy standby build should back off instead of retrying every tick.
Found by reading
468ce53during the celeris#639 fd audit. The leak is visible in the code: every step is cited below. Its rate under real load has not been measured.The leak
Worker.runinengine/iouring/worker.go:SO_REUSEPORTlisten socket and stores it:createListenSocket(815),w.listenFD = listenFD(820).ready <- fmt.Errorf("worker %d ring setup: …")(827) and returns without closinglistenFD.ready <- fmt.Errorf("worker %d initial submit: …")(913) and returns without closinglistenFDor the ring.The worker's
runhas already returned, so itsshutdownnever runs.Engine.Listen(engine.go:261-265) cancels the other workers, joins them, and returns the error. Nothing ever closes the failed worker's listen socket.Why this is worse than a leaked descriptor
The socket is still listening, in the port's
SO_REUSEPORTgroup. The kernel keeps hashing a share of new connections to it. Nobody accepts from it, so those clients complete the handshake into a backlog that is never drained, and they hang.The adaptive engine is where that group is shared.
adaptive/engine.go:685for io_uring) startsbuilt.Listenon the same port as the active epoll engine, then waits 5 s forAddr().performSwitchlogsaborting switch: lazy standby build failed; staying on current activeand returns (686-689).e.secondarystays nil. The eval ticker is 1 s (controller.go:95). Unlessctrl.evaluatebacks off by itself (not checked here), the next tick builds and fails again.e.mufor the full 5 s bind wait.The engine is left serving on epoll with a growing fraction of its incoming connections stranded. No error comes back to the caller; only a
Warnis logged.What would make a worker's ring setup fail after the tier probe passed
The tier probe (
engine.go~220-238) creates and closes a single test ring. Per-worker rings are created later, one per worker.The plausible causes are:
SQPOLLlimits.The CI shape (
RLIMIT_MEMLOCK8 MiB, see #533's history) is where ring creation is known to be tight. Not yet measured.Fix
listenFD(and, on the submit path, the ring) on every early return after it is created. Adeferthat runs unless the worker reachedready <- nilcovers all of them.Listen's error: count listeners on the port via/proc/net/tcp, or trybindwithoutSO_REUSEPORT. The test must fail on main.