Skip to content

fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) - #932

Merged
FumingPower3925 merged 2 commits into
mainfrom
fix/celeris-862-eventloop-wake-close-race
Oct 5, 2026
Merged

FumingPower3925 merged 2 commits into
mainfrom
fix/celeris-862-eventloop-wake-close-race

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Lane D1, PR 1 of 3. The three PRs all edit driver/internal/eventloop/loop_linux.go and are stacked in this order: this one (#862), then #933 (#842, which carries this commit), then #934 (#881, which carries both). Each targets main so that CI runs on it. In the later two, review only the last commit.

Rebased onto main 8476e8b (2026-10-05)

This PR's head is now 92f2913 (commits 1807309 and 92f2913), rebased with git rebase --onto 8476e8b df31adc from dce5a3c (commits 6c25e65 and dce5a3c). The rebase applied with no conflicts, and each commit's git patch-id --stable is unchanged. None of main's 14 commits since df31adc touches driver/internal/eventloop or internal/wakefd, and this PR changes no exported API (it touches only driver/internal/eventloop). The controls below ran at dce5a3c on df31adc. They were all re-run at 92f2913 on 8476e8b (linux/arm64 golang:1.27 container, -race, one process per observation), and every arm's tally matches round 1's: ff-main 0/60/0 (PASS/FAIL/SKIP over the three tests × 20 processes; 20 race reports), ff-main-hooked 0/60/0 (20), fix 60/0/0 (0), nc 0/60/0 (20), mut-MREG 20/40/0, mut-MREG2 20/40/0, mut-MWAKE 40/20/0 (20), mut-EPFDCHECK 40/20/0. Whole suites at 92f2913: the eventloop package 29/0/0, and ./driver/... ./internal/wakefd/ 471/0/0, both with 0 race reports (464 before; the 7 added are main's #859 tests). Scripts: bash evidence/lanes-20261003/D1/fix-r4/scripts/trees.sh 8476e8b 92f2913 2feac79, then bash evidence/lanes-20261003/D1/fix-r4/scripts/controls.sh 862 20261005T0208Z. Logs: evidence/lanes-20261003/D1/fix-r4/logs/862-20261005T0208Z/. The timing A/B under Cost ran at dce5a3c and was not re-run, because the rebase changed no file this PR touches.

Defect

The standalone driver event loop (driver/internal/eventloop) is what the redis, memcached and postgres drivers use when they run without an engine. Its worker owns two descriptors that callers on other goroutines use: the wakeup eventfd and the epoll fd. Loop.Close closes both. Two sites used one of those numbers outside the lock that shutdown closes it under:

  • wake (the defect filed in eventloop: a Write racing Loop.Close can write the closed eventfd's number in wake(), which another socket may hold by then (data race) #862). wake read w.eventFD with no lock, then wrote 8 bytes to it. Meanwhile shutdown closed the eventfd and stored -1 under w.mu. A Write that leaves bytes pending wakes the worker after it has released c.mu, and it checks w.closed only on entry. So a Write running alongside Loop.Close was a data race. It could also write 01 00 00 00 00 00 00 00 to the number after the close, into whatever had taken the number by then: a socket, a driver conn, or another eventfd.
  • RegisterConn (same family, found in the audit). It read epollFD under w.mu, released w.mu, and only then issued EPOLL_CTL_ADD. A shutdown in between closed the epoll fd, and the ADD then had two possible outcomes:
    • It landed on the epoll instance that had taken the number. The conn's events then went to another loop.
    • It failed on the closed number. RegisterConn returned an epoll_ctl error for a conn whose onClose(ErrLoopClosed) had already fired.

Mechanism (base df31adc)

Line numbers are at 56a6c1e, where round 0's controls ran. driver/internal/eventloop and internal/wakefd are byte-identical at 93bd88f, where the triage probe failed, at df31adc, where round 1's controls below ran, and at 28383e8 (current main).

  • wake, driver/internal/eventloop/loop_linux.go:276-283: the unlocked check is at :277 and the write at :282.
  • shutdown, :235-273: it takes w.mu.Lock at :245, closes the eventfd at :255 and stores -1 at :258.
  • enqueueFlush calls wake at :476. It is reached after c.mu is released from these callers:
    • Write at :417. Its w.closed check runs only on entry, at :378.
    • WriteAndPoll at :574.
    • WriteAndPollBusy at :764.
    • WriteAndPollMulti at :960.
  • Loop.Close (loop.go:74-95) wakes the workers at :83 and joins only the worker goroutines at :86 (l.wg.Wait). It then runs shutdown at :90.
  • RegisterConn, :290-321: epfd := w.epollFD is read under w.mu at :308. w.mu is released at :309, and the ADD on epfd runs at :311.

Fix

  • The eventfd is owned through internal/wakefd.WakeFD. This is the handle fix(engine): own the wakeup eventfd so no producer can write it after shutdown closed it (celeris#655) #666 introduced for the engines' twin of this defect (epoll/io_uring: a transplant can write the wakeup eventfd after shutdown closed it, into whatever reused that descriptor number (found by reading, not reproduced) #655). Signal and Close take the same lock, so a wake either completes before the close or writes nothing. wake is now w.wakeFD.Signal(), and shutdown calls w.wakeFD.Close() under w.mu as before. The worker goroutine reads the number once, at the top of run, with the lock-free FD(). That is safe because shutdown closes the eventfd only after Loop.Close has joined run. One behaviour changes with it (review round 2): WakeFD.Close discards the eventfd's close(2) error, so Loop.Close now returns only the epoll fd's close error, where shutdown used to return the eventfd's too. The package is internal, and its only caller of Loop.Close discards the error (_ = l.Close() in Release, registry.go:113).
  • RegisterConn issues the ADD under c.mu, after a check of c.closed. Every other epoll_ctl of a conn already follows this rule (flushLocked, setEvents). The conn is in the map by then, and a conn leaves the map only once it is marked closed. shutdown marks every conn in the map closed, each under its c.mu, before it closes the epoll fd. So the ADD either completes before the close or is not issued. An UnregisterConn of the same fd marks the conn closed before its EPOLL_CTL_DEL, so the ADD comes before that DEL or not at all.
    • A conn torn down between entering the map and the ADD (by shutdown, or by an UnregisterConn racing the registration) has already had its onClose. No ADD is issued, and the registration is reported as made, the same as for a conn torn down right after it.
    • A conn whose ADD fails is marked closed before it leaves the map. No teardown then fires onClose for a registration that was refused.
    • w.mu is not held across the ADD (review round 1). Round 0 held its read lock there. A sync.RWMutex makes new readers wait once a writer is queued, so a forget, a registration or a shutdown queued behind the ADD made the worker's lookups wait for the syscall as well. v1.6.0 polish: drivers #887 already tracks the measured cost of an epoll_ctl held under the exclusive w.mu (forget's DEL, for a reader on the same worker under register/unregister churn). This ADD adds none.
  • testHookBeforeAdd, a nil-by-default test hook, sits next to fix(eventloop): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#784) #843's testHookBeforeRead, in the unlocked window between the map and the ADD.

Deadlock check (RULE 10).

  • wakefd's mutex is a leaf: nothing is acquired while it is held.
  • wake's callers hold no lock, or hold the conn's recvMu when they run inside an onRecv callback. recvMu comes first in the documented order (recvMu, w.mu, c.mu, c.rmu).
  • RegisterConn takes w.mu (write) to insert the conn and releases it, then takes c.mu, then c.rmu; when the ADD fails it takes w.mu again with no other lock held.
  • shutdown takes w.mu (write), then c.mu/c.rmu through markClosed. No path takes w.mu while it holds c.mu.

Tests and controls

Three tests are added in wake_close_862_linux_test.go:

  • TestWriteRacingCloseNeverTouchesTheClosedEventfd862 runs 64 rounds. In each, 4 goroutines Write 1 byte to a conn whose send buffer is full, so every Write takes the pending path and wakes the worker, while Loop.Close runs. It has two oracles:
    • Under -race, the report of the unlocked eventfd read. The test's coverage rests on this report; CI's root step runs this package with -race.
    • A second, opportunistic oracle: two eventfds opened as soon as Close returns take the numbers Close freed, and neither may receive a write. A wake has to land in the few microseconds between the close and the reuse for it to fire, so it seldom does, with or without -race (the "wake into a freed number" column below).
  • TestRegisterConnRacingCloseNeverAddsToAClosedEpoll862 is deterministic. The hook runs Loop.Close to completion in the window, then the test opens epoll instances until one takes the closed epoll fd's number. The ADD must not land there and must not fail on the closed number. onClose and the returned error must agree. If no epoll instance takes the number, the test fails rather than skipping that check (review round 1).
  • TestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862 (review round 1) is deterministic. The hook runs UnregisterConn of the same fd to completion in the window. Once both calls have returned, the fd must not be in the worker's epoll set: the test probes the set with an EPOLL_CTL_ADD of its own, which fails with EEXIST if it is. RegisterConn must return nil, and onClose must fire once. Mutant EPFDCHECK (the ADD skipped only once shutdown has retired the epoll fd, instead of whenever the conn was torn down) passes the other two tests and fails this one.

Each arm runs the tests as 20 separate processes (-race, linux/arm64, golang:1.27, kernel 7.0.12-linuxkit; one process is one observation). The race detector reports a given race once per process, which is why -count inside one process would understate the rate.

arm tree expected race test P/F/S register/Close test P/F/S register/unregister test P/F/S processes (failed) race reports processes with a wake into a freed number ADD landed on a foreign epoll A left in the epoll set log
ff-main origin/main + the #862 tests + a hook declaration (the base never calls the hook) all three FAIL (the register tests: the hook never ran) 0/20/0 0/20/0 ("hook never ran" lines, both register tests: 40) 0/20/0 20 (20) 20 0 0 0 fix-r1/logs/862-r1b/ff-main.log
ff-main-hooked origin/main + the #862 tests, the hook placed at the base's EPOLL_CTL_ADD (after w.mu is released) all three FAIL, the register tests on the defect itself 0/20/0 0/20/0 0/20/0 20 (20) 20 1 20 20 fix-r1/logs/862-r1b/ff-main-hooked.log
fix #862 head all three PASS 20/0/0 20/0/0 20/0/0 20 (0) 0 0 0 0 fix-r1/logs/862-r1b/fix.log
nc NC: base loop_linux.go back in with cp (+ hook declaration) all three FAIL 0/20/0 0/20/0 ("hook never ran" lines, both register tests: 40) 0/20/0 20 (20) 20 0 0 0 fix-r1/logs/862-r1b/nc.log
mut-MREG 2nd control: fix, with the ADD on the epoll fd number read in the first section and no closed check (the base's shape) both register tests FAIL, race test PASSes 20/0/0 0/20/0 0/20/0 20 (20) 0 0 20 20 fix-r1/logs/862-r1b/mut-MREG.log
mut-MREG2 2nd control: fix, with the ADD on w.epollFD but no closed check both register tests FAIL, race test PASSes 20/0/0 0/20/0 0/20/0 20 (20) 0 0 0 20 fix-r1/logs/862-r1b/mut-MREG2.log
mut-MWAKE 2nd control: fix, with wake reading a plain eventFD field again race test FAILs, register tests PASS 0/20/0 20/0/0 20/0/0 20 (20) 20 0 0 0 fix-r1/logs/862-r1b/mut-MWAKE.log
mut-EPFDCHECK 2nd control: fix, with the ADD skipped only once the epoll fd is retired (w.epollFD >= 0 instead of !c.closed) register/unregister test FAILs, the other two PASS 20/0/0 20/0/0 0/20/0 20 (20) 0 0 0 20 fix-r1/logs/862-r1b/mut-EPFDCHECK.log
  • main-pkg: PASS=26 FAIL=3 SKIP=0 race reports=1; FAIL: ['TestRegisterConnRacingCloseNeverAddsToAClosedEpoll862', 'TestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862', 'TestWriteRacingCloseNeverTouchesTheClosedEventfd862']; SKIP: [] (fix-r1/logs/862-r1b/main-pkg.log)
  • fix-pkg: PASS=29 FAIL=0 SKIP=0 race reports=0; FAIL: []; SKIP: [] (fix-r1/logs/862-r1b/fix-pkg.log)
  • fix-drivers: PASS=464 FAIL=0 SKIP=0 race reports=0; FAIL: []; SKIP: [] (fix-r1/logs/862-r1b/fix-drivers.log)

Failure rate on main: under -race, the race test fails in 20 of 20 processes, and both register tests fail on the defect itself whenever the hook is at the base's ADD (ff-main-hooked). The race test's opportunistic oracle (a wake into a freed number) fired only in the processes counted in that column; its detection rests on -race. The second controls show that each test catches its own defect and only that one.

This PR's head is unchanged in round 2 (dce5a3c), so these are round 1's controls, which ran at this head. Scripts: bash evidence/lanes-20261003/D1/fix-r1/scripts/trees.sh df31adc dce5a3c 399072a a509a99 53994be (the export trees, built from round 1's later heads; the git worktree is never mutated), then bash evidence/lanes-20261003/D1/fix-r1/scripts/controls.sh 862 r1b. Table: python3 evidence/lanes-20261003/D1/fix-r1/scripts/table862.py evidence/lanes-20261003/D1/fix-r1/logs/862-r1b. Paths are relative to the maintainer's evidence root (~/.claude/projects/-Users-fuming-Documents-github-celeris-probatorium/).

Main moved during round 2. #935 (#859) merged into main as 28383e8 while this round ran. It changes driver/postgres and adds close-after-teardown tests to the three drivers; none of the commits since the base touches driver/internal/eventloop or internal/wakefd. The stack is behind main, and it merges with it cleanly (git merge-tree --write-tree 28383e8 a21d0ff). The merged tree passes with -race: the eventloop package 39/0/0 (PASS/FAIL/SKIP) with 0 race reports, and ./driver/... ./internal/wakefd/ 481/0/0 with 0 race reports, including #935's 7 #859 tests (evidence/lanes-20261003/D1/fix-r2/logs/merged-28383e8-a21d0ff/; script bash evidence/lanes-20261003/D1/fix-r2/scripts/merged-suite.sh 28383e8 a21d0ff).

Cost

Measured: no cost resolved. Measured 2026-10-05, 01:32Z to 01:57Z, with bash evidence/lanes-20261003/D1/fix-r1/scripts/bench.sh 10 10 df31adc dce5a3c aeeccf7 a21d0ff (the fix-r1 copy of the command, which differs from fix-r2/scripts/bench.sh only in its output directory; dce5a3c, aeeccf7 and a21d0ff are the current heads of #932, #933 and #934). It ran under the laptop TIMING lock, in one linux/arm64 golang:1.27 container (--cpus 4): 10 rounds, one process per arm per round, the arm order rotated each round, -benchtime 1s, and an A/A arm (df31adc's test binary run a second time as its own arm). benchstat: median ± 95% CI, n=10 per arm, ~ = not significant at α=0.05. Host: no game client ran. 12 of the 25 one-minute samples during the run were NOT-QUIET by the script's rule, each because of one desktop process (Orca Helper, 49% to 58% of one of the 8 cores); the other 13 were quiet. The A/A floor: on the micro benchmarks the A/A arm differs from df31adc by 1.2% to 1.4% in 2 of 5 rows at p=0.023 and p=0.029, so a shift under about 1.5% there is not resolved. BenchmarkReadWhileAnotherConnFlushes784's sec/op CIs are ±13% to ±384% per arm, and its gapped rows are bimodal in every arm, the A/A arm included (round trips cluster near 20 µs and near 26 µs), so only large shifts there are resolved. Output: evidence/lanes-20261003/D1/fix-r1/bench/20261005T013145Z/ (the per-arm *.txt, run.log with the test binaries' sha256, quiet-during.log, benchstat/). Analysis: evidence/lanes-20261003/D1/fix-r3/scripts/analyse.sh (benchstat) and fix-r3/scripts/tables.py (these tables).

#932 against df31adc. Each cell is the median ± CI. Each delta is against df31adc.

benchmark (sec/op) df31adc A/A (df31adc again) #932
HandleReadable784 808.1 ns ±3% 796.9 ns ±3% -1.39% (p=0.023) 802.8 ns ±4% ~ (p=0.393)
WriteAndPoll784/WriteAndPoll 1.986 µs ±2% 1.983 µs ±3% ~ (p=0.698) 2.012 µs ±3% ~ (p=0.382)
WriteAndPoll784/WriteAndPollBusy 1.996 µs ±2% 1.99 µs ±2% ~ (p=0.529) 1.98 µs ±3% ~ (p=0.197)
WriteAndPoll784/WriteAndPollMulti 1.984 µs ±2% 1.98 µs ±2% ~ (p=0.363) 1.99 µs ±2% ~ (p=0.781)
WritePending862 370.2 ns ±3% 365.7 ns ±2% -1.23% (p=0.029) 369.8 ns ±1% ~ (p=0.579)
RegisterChurn862/churn=0 19.04 µs ±30% 21.23 µs ±17% ~ (p=0.247) 19.22 µs ±31% ~ (p=0.529)
RegisterChurn862/churn=1 27.67 µs ±13% 27.32 µs ±12% ~ (p=0.912) 24.99 µs ±14% ~ (p=0.481)
RegisterChurn862/churn=4 101.8 µs ±17% 103 µs ±15% ~ (p=0.971) 97.52 µs ±16% ~ (p=0.796)
ReadWhileAnotherConnFlushes784/chunk=0/gap=0s 19.37 µs ±30% 21.73 µs ±16% ~ (p=0.529) 19.26 µs ±31% ~ (p=0.631)
ReadWhileAnotherConnFlushes784/chunk=4096/gap=0s 142.6 µs ±29% 164.3 µs ±44% ~ (p=0.529) 148.4 µs ±28% ~ (p=0.796)
ReadWhileAnotherConnFlushes784/chunk=65536/gap=0s 5.115 ms ±384% 9.139 ms ±259% ~ (p=0.739) 9.942 ms ±474% ~ (p=0.393)
ReadWhileAnotherConnFlushes784/chunk=524288/gap=0s 1.539 ms ±69% 1.028 ms ±131% ~ (p=0.579) 496.9 µs ±230% ~ (p=0.190)
ReadWhileAnotherConnFlushes784/chunk=262144/gap=1ms 21.83 µs ±22% 21.4 µs ±19% ~ (p=0.725) 19.93 µs ±13% ~ (p=0.393)
ReadWhileAnotherConnFlushes784/chunk=1048576/gap=5ms 19.84 µs ±28% 20.06 µs ±18% ~ (p=0.796) 19.42 µs ±16% ~ (p=0.218)
benchmark (allocs/op) df31adc A/A (df31adc again) #932
HandleReadable784 0 ±0% 0 ±0% ~ (p=1.000) 0 ±0% ~ (p=1.000)
WriteAndPoll784/WriteAndPoll 0 ±0% 0 ±0% ~ (p=1.000) 0 ±0% ~ (p=1.000)
WriteAndPoll784/WriteAndPollBusy 0 ±0% 0 ±0% ~ (p=1.000) 0 ±0% ~ (p=1.000)
WriteAndPoll784/WriteAndPollMulti 0 ±0% 0 ±0% ~ (p=1.000) 0 ±0% ~ (p=1.000)
WritePending862 0 ±0% 0 ±0% ~ (p=1.000) 0 ±0% ~ (p=1.000)
RegisterChurn862/churn=1 9.5 ±16% 9 ±11% ~ (p=0.750) 9 ±11% ~ (p=1.000)
RegisterChurn862/churn=4 33.5 ±16% 33 ±15% ~ (p=0.812) 32 ±16% ~ (p=0.838)
benchmark (pairs/s) df31adc A/A (df31adc again) #932
RegisterChurn862/churn=1 251.8k ±2% 249.8k ±2% ~ (p=0.684) 249.6k ±3% ~ (p=0.529)
RegisterChurn862/churn=4 301.9k ±2% 299.9k ±2% ~ (p=0.971) 301.2k ±2% ~ (p=0.796)

No row differs from df31adc at α=0.05 in sec/op or pairs/s, and allocs/op does not change on the paths this PR touches. BenchmarkWritePending862 (the Signal read lock): 369.8 ns against 370.2 ns (p=0.579). BenchmarkRegisterChurn862 (the ADD under c.mu): no shift at churn 0, 1 or 4 (p=0.48 to 0.80). The one significant row in any metric is chunk=524288 allocs/op, 5.5 down to 4.0 (p=0.023). That row's B/op is ±1004% in df31adc, so the shift is noise and not a saving.

What changes on the measured paths, without numbers: a Write that leaves bytes pending now takes WakeFD.Signal's read lock around the eventfd write(2) it already made (#666 measured that shape for the engines in internal/wakefd's BenchmarkSignal). RegisterConn takes w.mu once, as on main, and now issues its ADD under the new conn's c.mu, which is uncontended unless a call on the same fd or a teardown races the registration. The benchmarks committed for these paths are BenchmarkWritePending862 and BenchmarkRegisterChurn862; both run on main too.

Family audit

Every site where a caller on another goroutine uses one of the worker's own descriptors:

Site Descriptor Status
wake (enqueueFlush, Loop.Close) eventfd fixed (WakeFD)
RegisterConn ADD epoll fd fixed (c.mu, closed check)
flushLocked MODs (Write, WriteAndPoll* step 1, drainOne) epoll fd safe: under c.mu after the closed check; shutdown marks every conn closed before it closes the epoll fd
setEvents (the WriteAndPoll* mask and re-arm) epoll fd (read under w.mu) safe: same rule
forget DEL epoll fd safe: under w.mu with an epollFD >= 0 check
run (EpollWait, the eventfd drain) both safe: the worker goroutine is joined before shutdown
loop_other.go (the non-Linux fallback; it builds on darwin, and the drivers do not build on windows at all: #937) none no eventfd and no epoll; each reader holds its own dup

#655 is the engines' twin; #666 took the same approach there (the wakeup eventfd, and ctlMu for the epoll fd). The drivers do not own either descriptor.

Review round 1

  • MINOR (perf/API/security): RegisterConn held w.mu's read lock across the ADD, and its comment said that did not hold up the worker's lookups. That is not so once a writer is queued. The lock is dropped (see Fix); the controls, including mutants MREG and MREG2, were re-run at dce5a3c.
  • MINOR (correctness): the register-vs-unregister half of the ADD fix had no test. TestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862 is added, with mutant EPFDCHECK as its second control.
  • NIT: the wake test's comment now says that its coverage rests on -race and that the second oracle is opportunistic. A stronger non-race test is item "fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) #932 wake test" on v1.6.0 polish: CI and tests #890.
  • NIT: the register-vs-Close test fails instead of skipping its main check when no epoll instance takes the number.
  • MINOR (both): the cost. See Cost.

Review round 2

The correctness review approved; the perf/API/security review's MAJOR is #934's measurement, and its MINOR here is this PR's cost (see Cost). This PR's code is unchanged in round 2.

Fixes #862

@FumingPower3925 FumingPower3925 added this to the v1.6.0 milestone Oct 3, 2026
@FumingPower3925 FumingPower3925 added bug Something isn't working platform/linux Linux-specific (io_uring, epoll) area/driver Database/cache driver infrastructure labels Oct 3, 2026
@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: Repository: goceleris/celeris/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 33ae37ab-beb6-43bf-bd97-e3a4b9844c5a
📥 Commits

Reviewing files that changed from the base of the PR and between 8476e8b and 92f2913.

📒 Files selected for processing (4)
  • driver/internal/eventloop/loop_linux.go
  • driver/internal/eventloop/register_churn_862_bench_linux_test.go
  • driver/internal/eventloop/wake_close_862_linux_test.go
  • driver/internal/eventloop/write_pending_862_bench_linux_test.go
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 78.26087% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
driver/internal/eventloop/loop_linux.go 78.26% 5 Missing ⚠️

📢 Thoughts on this report? Let us know!

FumingPower3925 added a commit that referenced this pull request Oct 3, 2026
…d a RegisterConn racing UnregisterConn is tested (celeris#862)

Review round 1 of #932.

RegisterConn held w.mu's read lock across its EPOLL_CTL_ADD. A queued
writer of w.mu (a forget, a registration, shutdown) then made the worker's
lookups wait for the syscall, because a sync.RWMutex holds new readers back
once a writer waits. The read lock is not needed: c is in the map, a conn
leaves the map only once it is marked closed, and shutdown marks every conn
in the map closed, under its c.mu, before it closes the epoll fd. So c.mu
and the closed check alone keep the ADD off a closed or reused epoll fd
number, the rule flushLocked already relies on.

TestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862 covers the other
half of that check: an UnregisterConn of the same fd that runs in
RegisterConn's window must leave fd out of the epoll set. A check of the
epoll fd instead of c.closed passes both earlier tests and fails this one.

The register-vs-Close test now fails, rather than skipping its main check,
when no epoll instance takes the closed epoll fd's number, and the wake
test's comment says that its coverage rests on -race.
@FumingPower3925
FumingPower3925 force-pushed the fix/celeris-862-eventloop-wake-close-race branch from ec5095c to dce5a3c Compare October 3, 2026 13:43
…ses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862)

The standalone driver loop's worker published its wakeup eventfd's number in
a plain field. A Write that left bytes pending read it with no lock (wake,
via enqueueFlush, after c.mu is released) while shutdown closed the eventfd
and stored -1, so a Write running alongside Loop.Close was a data race, and
could write 8 bytes to the number after the close, into whatever had taken
it. The worker now holds the eventfd through internal/wakefd.WakeFD, the
handle the engines adopted for the same defect (celeris#655, #666): Signal
and Close share a lock, so a wake either completes before the close or
writes nothing.

RegisterConn had the same shape for the other descriptor a caller reaches:
it read epollFD under w.mu, released w.mu, and issued EPOLL_CTL_ADD after.
A shutdown in between closed the epoll fd, so the ADD went to the epoll
instance that had taken the number, or to a closed number (RegisterConn
then returned an epoll_ctl error for a conn whose onClose had already
fired). The ADD is now issued under w.mu's read lock and c.mu, after a
check that the conn has not been torn down, like every other epoll_ctl of a
conn; shutdown marks every conn closed and closes the epoll fd under the
write lock. A read lock, so a registration does not hold up the worker's
lookups. A conn whose ADD fails is marked closed before it leaves the map.

BenchmarkWritePending862 (the wake path) and BenchmarkRegisterChurn862 (a
conn's round trip while other goroutines register and unregister on its
worker) measure the cost; both run on the base too.
…d a RegisterConn racing UnregisterConn is tested (celeris#862)

Review round 1 of #932.

RegisterConn held w.mu's read lock across its EPOLL_CTL_ADD. A queued
writer of w.mu (a forget, a registration, shutdown) then made the worker's
lookups wait for the syscall, because a sync.RWMutex holds new readers back
once a writer waits. The read lock is not needed: c is in the map, a conn
leaves the map only once it is marked closed, and shutdown marks every conn
in the map closed, under its c.mu, before it closes the epoll fd. So c.mu
and the closed check alone keep the ADD off a closed or reused epoll fd
number, the rule flushLocked already relies on.

TestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862 covers the other
half of that check: an UnregisterConn of the same fd that runs in
RegisterConn's window must leave fd out of the epoll set. A check of the
epoll fd instead of c.closed passes both earlier tests and fails this one.

The register-vs-Close test now fails, rather than skipping its main check,
when no epoll instance takes the closed epoll fd's number, and the wake
test's comment says that its coverage rests on -race.
@FumingPower3925
FumingPower3925 force-pushed the fix/celeris-862-eventloop-wake-close-race branch from dce5a3c to 92f2913 Compare October 5, 2026 02:07
@FumingPower3925
FumingPower3925 marked this pull request as ready for review October 5, 2026 02:22
@FumingPower3925
FumingPower3925 merged commit 5260896 into main Oct 5, 2026
42 of 43 checks passed
@FumingPower3925
FumingPower3925 deleted the fix/celeris-862-eventloop-wake-close-race branch October 5, 2026 02:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/driver Database/cache driver infrastructure bug Something isn't working platform/linux Linux-specific (io_uring, epoll)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

eventloop: a Write racing Loop.Close can write the closed eventfd's number in wake(), which another socket may hold by then (data race)

1 participant