Skip to content

fix(engine): own the wakeup eventfd so no producer can write it after shutdown closed it (celeris#655) - #666

Merged
FumingPower3925 merged 4 commits into
mainfrom
fix/celeris-655-df1269c
Sep 16, 2026
Merged

FumingPower3925 merged 4 commits into
mainfrom
fix/celeris-655-df1269c

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Both engines published the number of their wakeup eventfd and let producers on other goroutines write it, then closed that descriptor at shutdown without resetting the field. Every producer that ran afterwards wrote 8 bytes into whatever the process had opened next. This is celeris#655, found by reading and now reproduced.

Review follow-up (round 2) widened this: epoll has a second descriptor with the same defect — its own epoll fd, which driver goroutines call epoll_ctl on. That is fixed here too, and the guard that was supposed to prove "no producer was missed" has been widened to the real class.

Closes #655

What was wrong

celeris#658 (PR #661) put the two adopt writers under the queue mutex shutdown takes before the close. That left every other producer, and on current main df1269c there were 24 raw unix.Write sites on those descriptors: 11 in engine/epoll, 12 in engine/iouring, 1 in internal/conn.

producer engine why it can outlive the close
addDriverAction (a driver's RegisterConn / UnregisterConn / Write) io_uring writes after releasing driverActionMu, no closed check. shutdownDrivers sets driverConns = nil and RegisterConn rebuilds the map from nil, so it always reaches the write
enqueueDetach io_uring called by the dispatch goroutine while an io_uring→epoll transplant drains — the switch-vs-shutdown window the issue describes
runAsyncHandler wakeups (5 sites) io_uring Worker.shutdown closes the eventfd before it joins the dispatch goroutines with asyncWG.Wait()
detached WS/SSE writeFn / PauseRecv / ResumeRecv both run on middleware goroutines asyncWG never tracked; the pause/resume pair has no detachClosed check at all
h2ShardedQueue.Enqueue both non-inline H2 streams run on globalH2Pool, which no engine joins, and CloseH2 → Manager.Close only cancels the streams. This queue has no engine lock of its own
driver epoll_ctl (RegisterConn / UnregisterConn / Write → flushDriverSendLocked / closeDriver) epoll the same defect on the epoll fd, not the eventfd. Loop.shutdown closes l.epollFD and never resets it; driver goroutines are joined by nothing. Deterministic for the same reason as row 1: shutdownDrivers nils the map, RegisterConn rebuilds it from nil and always reaches the epoll_ctl

Two things in the original report did not hold up against the code, and the fix is shaped accordingly:

  • epoll's runAsyncHandler and drainDetachQueue are not the hazard. Loop.shutdown joins its dispatch goroutines in phase 2, before the close in phase 3, and drainDetachQueue writes no eventfd. epoll's exposure is the producers asyncWG never tracked: the detached callbacks and the H2 queue. io_uring has the inverted, sharper ordering.
  • io_uring has no data race on the field. h2EventFD is assigned once before ready and never again, so a naive "set it to -1 in shutdown" would create one. The handle avoids that.

epoll does have a field race, and it is fixed here too: the lazy re-creation after a failed startup eventfd wrote l.eventFD on the loop thread while AdoptConn, the detached closures and switchToH2Local read it with no lock.

l.timerFD is loop-thread-only and is not a defect.

The fix

Generalise #661's rule — every write happens under a lock shutdown takes before it closes the fd — by moving the descriptor into one handle, internal/wakefd.WakeFD, that every producer signals through.

  • Signal and Close share an RWMutex, so a signal either completes before the close or is dropped, and Close waits for the signals already in flight.
  • Set lets the loop create its eventfd lazily and refuses once closed, which is how the caller learns the descriptor it just created is its own to close.
  • FD() serves the loop thread, the only goroutine allowed to touch the number directly. Producers never see a number, which is what removes epoll's field race.
  • New/Set force O_NONBLOCK on the descriptor, so Signal's write cannot park while holding the read lock Close waits behind.
  • A nil *WakeFD is a valid, permanently disabled handle — what a loop whose eventfd creation failed needs, and what the engine test fixtures that used to set the field to -1 now express by leaving it unset.

The same rule, applied to epoll's other shared descriptor: every driver-path epoll_ctl goes through Loop.driverEpollCtl, which holds a read lock that closeEpollFD takes as a writer before closing. A driver that reaches a loop after shutdown is refused with errLoopShutdown rather than operating on a recycled number. A side effect worth naming: RegisterConn after shutdown now fails instead of being silently accepted and never getting its onClose.

The coalescing is untouched: detachQPending.Swap and the H2 pending CAS still limit Signal to the queue's empty-to-non-empty edge.

Cost: measured, not asserted

The previous revision of this PR claimed "the added cost is one RLock/RUnlock beside a write(2) that was already there". That was wrong for FD(), and FD() is the call the epoll event loop makes for every event it dispatches, as the first branch of for i := range n:

if efd := l.wakeFD.FD(); fd == efd && efd >= 0 {

where main did a plain field load. It is fixed rather than excused: FD() is now a relaxed atomic load of a number written exactly twice in a loop's life (Set at startup, Close at shutdown), padded onto its own cache line so the producers' RLock in Signal never dirties the line the loop thread reads. The barrier stays where the defect is — on Signal.

internal/wakefd/wakefd_bench_test.go measures all four shapes, idle and with 4 goroutines signalling without pause (the WS-broadcast / H2 fan-out shape this engine is tuned for). golang:1.27, --cpus 4, linux/arm64, -benchtime=300ms -count=5:

FD() shape idle 4 producers signalling
plain field (main df1269c) 0.66 ns 0.70 ns
RWMutex RLock (this PR, as reviewed) 3.44 ns 18.16 ns
atomic, sharing the mutex's cache line 0.70 ns 0.84 ns
atomic, padded onto its own line (ships) 0.84 ns 0.92 ns

Median of 5, ns/op. The reviewed shape cost 5.2x idle and 26x under four
producers — 18.16 ns against 0.70 ns, paid once per epoll event, on the branch that
runs before every other one. The lock-free load puts that back within ~0.2 ns of
main's plain field.

Reported honestly: the padding is not a measurable win on this box (0.92 vs
0.84 ns contended, 0.84 vs 0.70 idle — both inside the run-to-run spread), because
each producer spends ~145 ns inside write(2) and so dirties the line rarely. It
is kept as insurance for a busier producer mix, costs one 60-byte hole per event
loop, and is recorded here as insurance rather than claimed as a speedup.

Signal — the producer side, which does keep the barrier — is the path where the original claim was true:

Signal ns/op
raw write(2) on the published number (main) 142.7
WakeFD.Signal (RLock + the same write(2)) 144.7

+2.0 ns, +1.4%, on a path that was already making a syscall — which is what the
original "one RLock/RUnlock beside a write(2)" sentence described, and the
only path it was true for.

Lock order

WakeFD.mu is a leaf: nothing is acquired while it is held. Signal is therefore legal under adoptQMu (epoll AdoptConn) and driverActionMu (io_uring addAdoptAction), and both keep taking those locks across the signal because they also publish the queue entry. Close runs on the loop thread holding no producer lock. #661's e.mu → wakeMu order is unchanged, and Signal is never called under wakeMu.

Loop.ctlMu (the new epoll_ctl guard) is a leaf on the same terms: held only across the epoll_ctl itself, and taken while holding driverMu (RegisterConn) or dc.mu (flushDriverSendLocked), never the reverse. closeEpollFD takes it as a writer at the very end of shutdown, holding nothing.

Worker.shutdown's ordering is deliberately not reordered: the conn-fd close before asyncWG.Wait() is protected by detachMu + asyncClosed, and moving the join would have fixed only two of the seven producers.

Evidence

All runs in a CI-shape container: golang:1.27, --cpus 4, -race, one container at a time. epoll and internal/conn at memlock 8 MiB, io_uring at 128 MiB. Tallies come only from --- PASS: / --- FAIL: lines. Every number below is reproducible from all-runs.sh and guard.sh in the run's log directory.

Each test drives a producer after shutdown has closed the descriptor, having first parked a stand-in on that exact descriptor number (F_DUPFD_CLOEXEC, which returns the lowest free number ≥ n, so it can never clobber a live fd). The stand-in is created before the close, or it would take the freed number itself. Every test also runs its producer on a live loop first and asserts it reaches the syscall — so a pass cannot mean the producer is simply inert.

Regression tests, -count=5 on the fix

test package producer class
TestRegisterConnAfterShutdownDoesNotWriteTheClosedWakeupFD engine/iouring addDriverAction
TestEnqueueDetachAfterShutdownDoesNotWriteTheClosedWakeupFD engine/iouring enqueueDetach
TestRunAsyncHandlerAfterShutdownDoesNotWriteTheClosedWakeupFD engine/iouring runAsyncHandler (new this round)
TestDetachedResumeRecvAfterShutdownDoesNotWriteTheClosedWakeupFD engine/iouring detached ResumeRecv
TestDetachedResumeRecvAfterShutdownDoesNotWriteTheClosedWakeupFD engine/epoll detached ResumeRecv
TestRegisterConnAfterShutdownDoesNotTouchTheClosedEpollFD engine/epoll driver epoll_ctl (new this round)
TestH2WriteQueueDoesNotSignalAClosedWakeupFD internal/conn H2 write queue

./engine/iouring PASS=20 FAIL=0 SKIP=0, ./engine/epoll + ./internal/conn PASS=15 FAIL=0 SKIP=0, 0 data races (4 tests × 5 and 3 tests × 5).

The epoll_ctl test's witness is not "an error was returned": a second epoll instance is parked on the closed descriptor's number, and the test asserts that instance did not silently gain the connection.

Negative controls — two, each targeting one barrier

Both neuters are applied to the fix and restored by cp, with the sha256 checked before, after and after restore.

A — remove the WakeFD barrier (delete closed = true / fd = -1 / num.Store(-1) from Close, leaving main's "close the descriptor, keep the number" semantics; the API is untouched so everything still compiles):

state sha256 of internal/wakefd/wakefd.go result
fix 7333ec73d905d230… all 7 PASS
barrier removed 8da378ecf31bf0f4… 6 FAIL — every eventfd test, with the byte-level witness
restored (cp) 7333ec73d905d230… all 7 PASS
RegisterConn wrote 01 00 00 00 00 00 00 00 into descriptor 6 after shutdown closed
the worker's wakeup eventfd; whatever now holds that number receives those bytes (celeris#655)

B — remove the epoll_ctl guard (restore main's semantics exactly: unguarded epoll_ctl, and a close that leaves the number in the field — otherwise EpollCtl(-1) would fail with EBADF and the test would pass for the wrong reason):

state sha256 of engine/epoll/driver.go result
fix 45338a015ffeef6b… epoll_ctl test PASS
guard removed e350c615ab34be7c… FAIL
restored (cp) 45338a015ffeef6b… PASS
RegisterConn added fd 10 to the epoll instance that now holds descriptor 6 after
shutdown closed the loop's epoll fd; that instance is now watching a connection
nobody registered with it (celeris#655)

The two controls are specific: A fails the 6 eventfd tests and leaves the epoll_ctl test passing; B fails the epoll_ctl test and leaves the eventfd tests passing. Neither neuter is a blanket break.

Static guard, widened (review MAJOR 2)

The old guard grepped only for unix.Write to a wakeup eventfd, so it could not have caught the epoll_ctl defect it was being cited to rule out. The class it now covers: a syscall from a goroutine the loop does not join, on a descriptor the loop closes at shutdown and whose number it does not reset.

check main df1269c fix
raw write(2) on a wakeup eventfd (engine/epoll + engine/iouring + internal/conn) 24 (epoll 11, io_uring 12, conn 1) 0 — WakeFD.Signal is the only one left
raw epoll_ctl on l.epollFD in the driver path (engine/epoll/driver.go) 6 1 — driverEpollCtl, the single guarded chokepoint
raw unix.Write / EpollCtl in engine/iouring/driver.go (io_uring's driver path enqueues; the worker does the ring work) 0 0
engine fields still publishing a wakeup descriptor number 2 0

guard.sh also enumerates every remaining raw epoll_ctl on l.epollFD with its enclosing function, so the "all of these are loop-thread" claim is auditable rather than asserted. The 13 remaining sites are in run, acceptAll, drainRead, hijackConn, initProtocol, drainDetachQueue, armEpollOut, disarmEpollOut, closeConn, attachAdoptedFD and detachFromEpoll — all loop-thread, except hijackConn, which in async mode runs on a dispatch goroutine that epoll joins in shutdown phase 2, before the phase-3 close. flushWrites (the one path a detached middleware goroutine reaches) touches no epoll fd.

Packages on the fix (-race -count=1 -v)

package memlock PASS FAIL SKIP data races
./engine/iouring 128 MiB 127 0 0 0
./engine/epoll + ./internal/conn + ./internal/wakefd 8 MiB 100 0 5 0

ok engine/iouring 109.890s | ok engine/epoll 20.608s | ok internal/conn 1.322s | ok internal/wakefd 1.007s

The previous revision recorded 126 and 97. The +1 and +3 are exactly the tests added
this round: runAsyncHandler in io_uring, and the epoll epoll_ctl regression test
plus TestConcurrentSetAndSignal and TestNewForcesNonBlocking in internal/wakefd.

The five SKIPs are the tests' own pre-existing gates, the same ones #661 recorded: TestWriteBufBackpressureClosesSlowConsumer needs GOTEST_BACKPRESSURE=1, and the four TestSendfile* tests build a 1-worker engine that config validation rejects. None is caused by this change, and no test was removed or newly skipped.

#661's own tests still pass: TestAdoptConnWakesASuspendedWorker, TestAdoptConnWakesASuspendedLoop and TestShutdownClosesAQueuedAdoption (both engines).

gofmt clean, go vet ./... clean on linux/arm64 and linux/amd64, golangci-lint run on the four packages: 0 issues..

Review points answered

review point disposition
MAJOR — unmeasured RWMutex in the epoll hot path, false cost claim Fixed and measured. FD() is lock-free and cache-line isolated; the claim is replaced by the benchmark table above
MAJOR — static guard does not cover the class; epoll driver path still defective Fixed. Guard widened and now enumerating; the epoll_ctl defect is fixed with its own test and negative control
minor — epoll field race (E4) asserted fixed with no test Fixed. TestConcurrentSetAndSignal races Set against Signal under -race, which is the E4 shape (TestConcurrentSignalAndClose only covered Signal vs Close)
minor — 3 of 7 producer classes untested, runAsyncHandler among them Fixed for runAsyncHandler (new test; the handler-panic teardown is the one of its five wake sites reachable without a live ring). Disputed for the guarded writeFn, in both engines: after shutdown it cannot reach its wakeup write at all, because the closure's first act is mu.Lock(); if cs.detachClosed { return } and shutdown sets detachClosed under that same mutex before closing the eventfd. A post-shutdown test would pass on unfixed main too — the producer is inert, not barred. Its real exposure is the in-flight window (passed the check, released the mutex, inside Signal when Close runs), which is a race rather than a sequence a deterministic test can pin; that window is what Close's wait for in-flight signals shuts, covered by TestConcurrentSignalAndClose. The reasoning is recorded beside the tests, not only here
minor — committed comments cite main's line numbers this diff shifted Fixed. They name functions and shutdown phases now, so they cannot go stale
minor — Signal's "bounded time" is a caller contract the type does not enforce Fixed. New/Set force O_NONBLOCK; TestNewForcesNonBlocking hands the type a deliberately blocking, already-full pipe and asserts both the flag and that Close still returns
nit — prepareH2Poll is the only FD() consumer with no >= 0 guard Fixed. Guarded before GetSQE, so a disabled handle consumes no SQE either
nit — PR has no labels and no milestone Fixed. bug, engine/epoll, engine/iouring, milestone v1.6.0 — matching #661 and #663

Out of scope

driver/internal/eventloop/loop_linux.go has the same bug class — wake() reads w.eventFD without w.mu while Close closes it and sets -1 under w.mu — but it is the driver's own event loop, not the engine wakeup path celeris#655 is about. Worth a separate issue.

What this change does not alter: a producer that runs after shutdown still enqueues onto a queue nobody will drain (driver actions, detach entries, H2 frames). That costs memory until the engine is collected; no descriptor is involved.

Test plan

  • Unit tests added (internal/wakefd, one regression test per producer class in both engines, the H2 queue, and the epoll driver epoll_ctl path)
  • Benchmarks added for the hot-path cost (internal/wakefd/wakefd_bench_test.go)
  • go vet ./... + golangci-lint run + gofmt
  • Tested on Linux (engine change): epoll, io_uring, and the shared H2 queue

Tested on: [ ] std [x] epoll [x] io_uring — [ ] amd64 [x] arm64

Every run above executed on linux/arm64 containers. linux/amd64 and linux/arm64 are both
compile- and vet-checked (GOOS=linux GOARCH=... go vet ./...), but only linux/arm64 was executed
here; the other arch needs a cluster run.

Release notes

  • Breaking change
  • bug

… shutdown closed it (celeris#655)

Both engines published the NUMBER of their wakeup eventfd and let producers
on other goroutines write it, then closed that descriptor at shutdown without
resetting the field. Every producer that ran afterwards wrote 8 bytes into
whatever the process had opened next: a spurious wakeup on another eventfd,
bytes injected into a client socket, silent corruption of a file. Nothing
reported any of it.

celeris#658 put the two ADOPT writers under the queue mutex shutdown takes
before the close. Seven producers were left: io_uring addDriverAction (a
driver's RegisterConn / UnregisterConn / Write), io_uring enqueueDetach, the
runAsyncHandler wakeups in both engines, the detached WS/SSE write and
PauseRecv / ResumeRecv closures in both engines, and the H2 write queue,
which has no engine lock at all and is drained by pool goroutines no engine
joins. io_uring is the sharper of the two: Worker.shutdown closes the eventfd
BEFORE it joins its dispatch goroutines.

The fix generalises #658's rule — every write happens under a lock shutdown
takes before it closes the fd — by moving the descriptor into one handle,
internal/wakefd.WakeFD, that every producer signals through. Signal and Close
share an RWMutex, so a signal either completes before the close or is dropped;
Close waits for the signals already in flight. The mutex is a leaf, so Signal
stays legal under adoptQMu and driverActionMu, and the coalescing that limits
it to the queue's empty-to-non-empty edge is untouched.

Holding the handle instead of the number also removes epoll's data race on
Loop.eventFD: the lazy re-creation after a failed startup eventfd wrote the
field on the loop thread while AdoptConn, the detached closures and
switchToH2Local read it unsynchronised. Set now takes the same lock Signal
reads under, and refuses once the handle is closed so the caller closes the
descriptor it just created.

A nil *WakeFD is a valid, permanently disabled handle, which is what a loop
whose eventfd could not be created needs — and what the engine test fixtures
that used to set the field to -1 now express by leaving it unset.

Raw eventfd writes in engine/epoll, engine/iouring and internal/conn: 24
before, 0 after. WakeFD.Signal is the only one left.

internal/wakefd carries its own tests for the barrier, for the lazy-Set
refusal that hands the descriptor back to the caller, for the nil handle,
and for a -race Signal-vs-Close storm.
@FumingPower3925 FumingPower3925 added this to the v1.6.0 milestone Sep 16, 2026
@FumingPower3925 FumingPower3925 added bug Something isn't working engine/epoll Epoll engine specifics engine/iouring io_uring engine specifics labels Sep 16, 2026
… the driver's epoll_ctl the same way (celeris#655)

Review follow-up on #666.

WakeFD.FD is no longer a lock. The epoll event loop calls it for every
event it dispatches, as the FIRST branch of `for i := range n`, where main
did a plain field load; routing that through the producers' RWMutex put two
atomic RMWs on the engine's hottest branch and shared a cache line with
every goroutine signalling. FD is now a relaxed atomic load of a number
written exactly twice in a loop's life (Set at startup, Close at shutdown),
padded off the mutex's line so the producers never dirty it. Signal keeps
the barrier — that is where the fix lives — and the PR body's cost claim is
replaced by a measurement (internal/wakefd/wakefd_bench_test.go benchmarks
all four shapes: plain field, RWMutex, atomic sharing the line, atomic
padded).

The static guard was too narrow, and it hid a live defect. The class is not
"a raw write to the wakeup eventfd", it is "a syscall from a goroutine the
loop does not join, on a descriptor the loop closes at shutdown and whose
number it does not reset". epoll has a second such descriptor: its own
epoll fd. Loop.shutdown closed it and left the number in place, while a
driver goroutine's RegisterConn / UnregisterConn / Write issues epoll_ctl
on it — deterministically after shutdown, since shutdownDrivers nils
driverConns and RegisterConn rebuilds the map from nil. Every driver-path
epoll_ctl now goes through driverEpollCtl, which holds a read lock
closeEpollFD takes as a writer before closing, so the operation either
completes first or is refused with errLoopShutdown. A registration after
shutdown now also fails instead of being silently accepted and never
getting its onClose.

Also from the review:

- internal/wakefd: New and Set force O_NONBLOCK on the descriptor, so
  Signal's bounded-time guarantee — which Close depends on, since it waits
  behind the signals in flight — is a property of the type rather than a
  convention every caller has to remember.
- prepareH2Poll: guard efd < 0 before arming, like every other FD()
  consumer. The handle can now report -1 where the raw field could not,
  and encodeUserData would have smeared it across the fd bits.
- Tests: the epoll field race (Set vs Signal) and io_uring's
  runAsyncHandler wake now each have a regression test, and the new
  epoll_ctl defect has one whose witness is a second epoll instance parked
  on the recycled number. The guarded writeFn is deliberately left untested
  and the reason is recorded next to the tests: shutdown sets detachClosed
  under the same mutex the closure takes first, so a post-shutdown test
  would pass on unfixed main too.
- Committed comments no longer cite line numbers this diff shifted; they
  name the functions instead.
…eFailedInit

celeris#664 added releaseFailedInit, which closed w.h2EventFD directly.
This branch replaces that field with a wakefd.WakeFD. The two merged
cleanly because they touch different lines, and the result did not
compile (worker.go:5267-5269, w.h2EventFD undefined).

Use w.wakeFD.Close(), which is what shutdown already calls on this
branch: it waits behind the signals in flight, retires the number before
the close, and is idempotent.
@FumingPower3925
FumingPower3925 merged commit 985a386 into main Sep 16, 2026
12 checks passed
@FumingPower3925
FumingPower3925 deleted the fix/celeris-655-df1269c branch September 16, 2026 07:34
FumingPower3925 added a commit that referenced this pull request Oct 3, 2026
…ses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862)

The standalone driver loop's worker published its wakeup eventfd's number in
a plain field. A Write that left bytes pending read it with no lock (wake,
via enqueueFlush, after c.mu is released) while shutdown closed the eventfd
and stored -1, so a Write running alongside Loop.Close was a data race, and
could write 8 bytes to the number after the close, into whatever had taken
it. The worker now holds the eventfd through internal/wakefd.WakeFD, the
handle the engines adopted for the same defect (celeris#655, #666): Signal
and Close share a lock, so a wake either completes before the close or
writes nothing.

RegisterConn had the same shape for the other descriptor a caller reaches:
it read epollFD under w.mu, released w.mu, and issued EPOLL_CTL_ADD after.
A shutdown in between closed the epoll fd, so the ADD went to the epoll
instance that had taken the number, or to a closed number (RegisterConn
then returned an epoll_ctl error for a conn whose onClose had already
fired). The ADD is now issued under w.mu's read lock and c.mu, after a
check that the conn has not been torn down, like every other epoll_ctl of a
conn; shutdown marks every conn closed and closes the epoll fd under the
write lock. A read lock, so a registration does not hold up the worker's
lookups. A conn whose ADD fails is marked closed before it leaves the map.

BenchmarkWritePending862 (the wake path) and BenchmarkRegisterChurn862 (a
conn's round trip while other goroutines register and unregister on its
worker) measure the cost; both run on the base too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working engine/epoll Epoll engine specifics engine/iouring io_uring engine specifics

Projects

None yet

1 participant