Split out of celeris#654 / PR #665, which found it while reading the same path. Not measured on epoll yet — read at df1269c.
Mechanism
With AsyncHandlers: true, runAsyncHandler holds cs.detachMu across the entire ProcessH1 call, i.e. for the whole user handler. Two loop-thread sites then block on that mutex:
checkTimeouts (engine/epoll/loop.go:2575) → closeConn (:2651) → cs.detachMu.Lock() (:2673). Reached whenever a deadline expires on a conn with a handler in flight — and nothing refreshes cs.lastActivity while a handler runs, so with the default ReadTimeout of 60 s any handler that outlives its read deadline parks the loop thread until it returns.
- The dirty-flush pass (
:605) takes cs.detachMu for every dirty conn on every iteration, so a slow handler holding the lock on a conn with pending bytes stalls the loop there too.
While the loop thread is parked, every other connection on that loop waits: no epoll_wait, no accepts, no flushes.
Relation to the closed io_uring twin
celeris#593 is the same family on io_uring and is CLOSED, fixed in bb4d774 ("fix(iouring): TryLock the timeout sweep's h1State snapshot so a slow async handler no longer pins the worker", PR #604): the sweep now goes through snapshotH1Deadlines, which uses TryLock and skips the conn for one pass rather than blocking the LockOSThread'd worker.
Two differences worth stating, because they are why this was not just the same bug:
It does still change when a timeout closes a conn — from "at the deadline" to "at or after the deadline, once the handler yields" — and on epoll it also has to decide whether to defer the close to the dispatch goroutine's detachQueue hand-off (when cs.asyncRun && !cs.asyncParked) rather than skip a pass. That is why it is its own issue rather than a rider on #665.
How to measure it
The celeris#589 rig (negctrl_async, its ANOMALY589 lines) reported 0 stalled samples on epoll in 40/40 runs — because that rig's 300 ms handler never outlives a timeout, so site 1 is never entered. To exercise it, set ReadTimeout below the handler duration (e.g. ReadTimeout: 100ms with a 500 ms async handler), keep a second fast keep-alive conn on the same loop via SO_REUSEPORT co-location, and sample the fast conn's latency: it should show a stall of roughly (handler duration − time already elapsed) once per reap, and none after a fix.
Split out of celeris#654 / PR #665, which found it while reading the same path. Not measured on epoll yet — read at
df1269c.Mechanism
With
AsyncHandlers: true,runAsyncHandlerholdscs.detachMuacross the entireProcessH1call, i.e. for the whole user handler. Two loop-thread sites then block on that mutex:checkTimeouts(engine/epoll/loop.go:2575) →closeConn(:2651) →cs.detachMu.Lock()(:2673). Reached whenever a deadline expires on a conn with a handler in flight — and nothing refreshescs.lastActivitywhile a handler runs, so with the defaultReadTimeoutof 60 s any handler that outlives its read deadline parks the loop thread until it returns.:605) takescs.detachMufor every dirty conn on every iteration, so a slow handler holding the lock on a conn with pending bytes stalls the loop there too.While the loop thread is parked, every other connection on that loop waits: no epoll_wait, no accepts, no flushes.
Relation to the closed io_uring twin
celeris#593 is the same family on io_uring and is CLOSED, fixed in
bb4d774("fix(iouring): TryLock the timeout sweep's h1State snapshot so a slow async handler no longer pins the worker", PR #604): the sweep now goes throughsnapshotH1Deadlines, which usesTryLockand skips the conn for one pass rather than blocking theLockOSThread'd worker.Two differences worth stating, because they are why this was not just the same bug:
h1Statewithout the lock — that contrast is in io_uring: checkTimeouts locks cs.detachMu for every live conn while runAsyncHandler holds it across ProcessH1, so a slow async handler pins the whole worker #593's own body — so on epoll only a conn that is actually being reaped blocks the loop. Narrower trigger, same consequence once it fires.It does still change when a timeout closes a conn — from "at the deadline" to "at or after the deadline, once the handler yields" — and on epoll it also has to decide whether to defer the close to the dispatch goroutine's
detachQueuehand-off (whencs.asyncRun && !cs.asyncParked) rather than skip a pass. That is why it is its own issue rather than a rider on #665.How to measure it
The celeris#589 rig (
negctrl_async, itsANOMALY589lines) reported 0 stalled samples on epoll in 40/40 runs — because that rig's 300 ms handler never outlives a timeout, so site 1 is never entered. To exercise it, setReadTimeoutbelow the handler duration (e.g.ReadTimeout: 100mswith a 500 ms async handler), keep a second fast keep-alive conn on the same loop via SO_REUSEPORT co-location, and sample the fast conn's latency: it should show a stall of roughly (handler duration − time already elapsed) once per reap, and none after a fix.