Found by reading 468ce53 during the celeris#639 fd audit. Not reproduced yet.
The shape
Both engines let another goroutine write the loop's wakeup eventfd, and both close that eventfd at shutdown without resetting the field.
epoll
Engine.AdoptConn (engine/epoll/adopt.go) runs on the transplanting goroutine: if l.eventFD >= 0 { unix.Write(l.eventFD, val[:]) }.
Loop.shutdown does unix.Close(l.eventFD) (loop.go:2809-2810) and leaves the number in the field.
- The loop thread also re-assigns
l.eventFD lazily (loop.go:1637, loop.go:2249). The unsynchronised read in AdoptConn is therefore also a data race with those writes.
- Callers: the io_uring → epoll handoff (
engine/iouring/transplant_source.go:143, :281).
io_uring
Worker.enqueueDetach (transplant_source.go:197-205) writes w.h2EventFD if it is >= 0.
Worker.shutdown closes it (worker.go:5114-5115) without a reset.
Why it would matter
A write that lands after the close goes to whatever now holds that number. That is 8 bytes (01 00 00 00 00 00 00 00), and what they do depends on the new owner:
| New owner of the number |
Effect of the late write |
| Another eventfd |
A spurious wakeup |
| A client socket |
The bytes are injected into its response stream |
| A file or pipe |
Silent corruption |
Nothing reports any of it.
When it could happen
A transplant handoff racing the target sub-engine's shutdown. Examples: a switch that is still draining connections when Shutdown cancels the contexts, or a lazily built standby that is torn down while the other engine is still handing connections to it. I have not measured whether that window is reachable in practice.
Direction for a fix
Choose one:
- Guard the field with the same lock producers take, and set it to -1 under that lock before closing.
- Or keep the eventfd open until every producer has been joined.
A -race test that transplants into an engine while shutting it down should expose the data race even when no fd reuse happens.
Found by reading
468ce53during the celeris#639 fd audit. Not reproduced yet.The shape
Both engines let another goroutine write the loop's wakeup eventfd, and both close that eventfd at shutdown without resetting the field.
epoll
Engine.AdoptConn(engine/epoll/adopt.go) runs on the transplanting goroutine:if l.eventFD >= 0 { unix.Write(l.eventFD, val[:]) }.Loop.shutdowndoesunix.Close(l.eventFD)(loop.go:2809-2810) and leaves the number in the field.l.eventFDlazily (loop.go:1637,loop.go:2249). The unsynchronised read inAdoptConnis therefore also a data race with those writes.engine/iouring/transplant_source.go:143,:281).io_uring
Worker.enqueueDetach(transplant_source.go:197-205) writesw.h2EventFDif it is >= 0.Worker.shutdowncloses it (worker.go:5114-5115) without a reset.Why it would matter
A write that lands after the close goes to whatever now holds that number. That is 8 bytes (
01 00 00 00 00 00 00 00), and what they do depends on the new owner:Nothing reports any of it.
When it could happen
A transplant handoff racing the target sub-engine's shutdown. Examples: a switch that is still draining connections when
Shutdowncancels the contexts, or a lazily built standby that is torn down while the other engine is still handing connections to it. I have not measured whether that window is reachable in practice.Direction for a fix
Choose one:
A
-racetest that transplants into an engine while shutting it down should expose the data race even when no fd reuse happens.