Skip to content

fix(websocket): apply pause/resume under the lock that decides them, and lift a stale pause (celeris#667, celeris#672) - #671

Merged
FumingPower3925 merged 16 commits into
mainfrom
fix/celeris-667-6136554
Sep 27, 2026
Merged

FumingPower3925 merged 16 commits into
mainfrom
fix/celeris-667-6136554

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR fixes celeris#667 and celeris#672 together, because the measurement says they must ship together: the #667 fix alone makes #672's wedge more likely.

  • celeris#667: chanReader decided pause/resume under pausedMu and applied the decision to the engine after releasing it, so the appender and the drainer could reach the engine in the opposite order from the one in which they decided. Both callbacks now run with pausedMu held.
  • celeris#672: requestPause applies a decision Append took from a depth snapshot before pausedMu was held. If the handler drained to empty in between, the pause landed on an empty buffer and nothing could ever lift it. requestPause now re-checks the watermark under the same lock after applying the pause, and resumes at once if the pause is already stale (arm C of the earlier A/B, code-identical to the measured variant).
  • pausedMu is released by defer, so a callback that panics cannot leave it held.
  • SetPauser is ordered with the engine worker (round 3): it writes the callbacks under pausedMu, and requestPause reads them under it. That closes a pre-existing data race, which the middleware/websocket: a pause decided on a stale depth snapshot wedges the connection permanently when the handler drains to empty first (distinct from #667) #672 re-check had widened.
  • The lock order is re-derived for celeris#666's RWMutex (round 3), and a forced-interleaving test shows no WakeFD writer can wait on pausedMu.

Closes #667
Closes #672
Refs #705

celeris#705 is not fixed here. It is the stranded spill: a chunk spilled after the handler drained the channel is never promoted. It predates this PR, and mutant M5 survives because of it. It will be fixed on top of this PR.

Measured end to end on real sockets at 371c76a (pre-registered; 192 rounds, one container each; main, an A/A floor, the #667 fix alone and this PR; the CI shape at memlock 8 MiB and at 128 MiB; the backpressure and echo tests under -race, plus the earlier A/B's echo rig):

Measured at the unit level (round 2, unchanged): the #672 wedge rate was 0/640 connections at cap256 and 0/640 at cap16, against main's 560/640 and 515/640. For the #667 fix alone it was x3.25 and x14 main's per-edge rate.

What this revision changes (round 5)

No code change and no new measurement. The head is still 371c76a, and its CI (run 36284528042) is unchanged. This round corrects the reporting (the re-review of round 4, two MINOR points):

  • The tail is reported in full. The declared pooled p99.9 op time is now in the body for every cell, with the reason it jumps: at 14,400 ops it is the 15th slowest op. "No more frequent" and "does not show up in the tail" are replaced: the tail effect is not resolved at this n, and a small tail cost in epoll is possible. The re-review's 2 ms count is reported, labelled post hoc. The round-4 reply is corrected in place.
  • Every number quoted now comes from a saved script:
    • the stall counts and their Fisher tests: round5/51-STALLS.txt;
    • the tail detail: 52-TAIL.txt;
    • the replication's power: 53-MDE-CONFIRM.txt. This corrects 12-CONFIRM.md, which said the replication detects +2.3% with about 80% power. The saved computation gives 0.93.
  • The Reproducing inventory is corrected. It now says which script made which number, and which driver ran which rounds.

What this revision changes (round 4)

No code change. The head is still 371c76a, and its CI (run 36284528042) is unchanged. This round adds:

  • The throughput A/B (review MAJOR 2), run as pre-registered, with the load gate amended in a hashed addendum before any timing round. It adds a pre-registered replication of the one borderline cell. See Throughput below.
  • Worker-stall and tail numbers, as a declared secondary deviation (the review's suggestion).
  • Corrected wording (review MINOR). Main's code stalled once in round 3, in the A/A arm, so "main 0/48" now reads "main's code 1/96". The stall p-values are labelled post hoc, and the pre-registered contrast stands next to them. The round-3 reply is corrected in place.

What this revision changes (round 3)

Fast-forward from d3472e7 to 371c76a:

commit what
9122766 test: TestChanReaderSetPauserOrderedWithAppend, failing-first (review NIT: the SetPauser race)
0a3492d fix: SetPauser writes, and requestPause reads, the callbacks under pausedMu
651d895 test: TestChanReaderWakeFDWritersNeverWaitOnPausedMu, and the lock-order note re-derived (review MINOR: celeris#666)
69bcd1c test: the wedge harness header builds the #667-only arm from 3b6c500, the arm actually measured (review NIT)
371c76a ci: the interlock requires six tests and ten subtests

The rest of the review is answered in this body:

What the rebase changed

Rebased from d4bc70b onto main 9f4d89b (7 commits: #666, #676, #677, #678, #680, #681, #687). None of them touches middleware/websocket; git range-diff shows the three original commits patch-identical (=), engineread.go at the #667 fix commit is still sha256 2b64141c…, and the only file both sides changed is ci.yml, where main added the celeris#657 steps above this PR's step without a conflict. Build, go vet and go test -c of ./middleware/websocket, ./engine/epoll and ./engine/iouring pass for linux/amd64 and linux/arm64; gofmt and golangci-lint (v2.13) report nothing.

What the rebase did move is the engine's PauseRecv/ResumeRecv closures (review point): the comments cited engine/iouring/worker.go:2200-2231 and engine/epoll/loop.go:1672-1703, which are now 2342/2356 and 1770/1784. The comments now name the closures instead of lines, so the next engine change cannot stale them again.

One commit in that range does bear on this PR: celeris#666 made the engines' pause/resume callbacks wake the loop through wakefd.WakeFD.Signal, which takes an RWMutex read lock, and those callbacks run under pausedMu. The round-2 lock-order note predated it; it is re-derived below.

celeris#672: mechanism and fix

Append sends a chunk, then decides from len(r.ch) >= highWater that the engine should pause, and requestPause applies that decision after taking pausedMu. If the handler drains the channel to empty in that gap, each of its resume checks runs while pausedState is still false and does nothing; then the pause lands on an empty buffer. Read re-evaluates the resume only after a successful dequeue, and none can happen again: the engine is paused, so it delivers nothing, and nothing is buffered.

The fix, at the end of requestPause, under the lock that applied the pause:

if r.resume != nil && len(r.ch) <= r.lowWater && !r.hasSpill() {
	r.pausedState = false
	r.resume()
}

Once pausedState is true under the lock, every later dequeue is followed by a resume check in Read that sees it, so the pause cannot go stale again after this critical section ends.

SetPauser is now ordered with the engine worker (review point)

tryEngineUpgrade registers Append as the data sink before Detach, and calls SetPauser only after Detach and after the 101 is written, on the goroutine running the upgrade. On an async-mode connection that goroutine is not the engine worker, and from the moment the 101 is on the wire the worker may be appending. An Append that crosses highWater read r.pause in requestPause without a lock, and the #672 re-check added a read of r.resume there. Nothing ordered those reads with SetPauser's writes: detachQMu, the only lock the upgrade shares with the worker, is released before SetPauser runs. The race predates this PR (main reads r.pause the same way); the re-check widened it by one field.

The fix is local (0a3492d). SetPauser writes both callbacks under pausedMu, and requestPause reads r.pause under it: the nil check moves below the Lock, and the re-check was already under it. With the callbacks installed, requestPause took pausedMu anyway, so nothing changes in steady state. Read's reads of r.resume stay unlocked, because they run on the handler goroutine, which the upgrade starts after SetPauser returns.

TestChanReaderSetPauserOrderedWithAppend (9122766, one commit before the fix) starts the two sides together, with nothing synchronising them. Under -race it fails on the commit before the fix at both reads (engineread.go:252 for r.pause and :283 for r.resume, each against SetPauser), and on main's engineread.go at the r.pause read. It passes on the fix. Without -race it skips rather than passing vacuously; CI runs it under -race. Mutants M10 (unlocked SetPauser) and M11 (nil check back above the Lock) are both killed by it.

Lock order after celeris#666 (review point)

The previous revision's safety case for holding pausedMu across the engine callbacks predated celeris#666, which the rebase brought under this branch. #666 made the callbacks wake the loop through wakefd.WakeFD.Signal, and Signal takes a sync.RWMutex read lock. So pausedMu now nests a read lock of an RWMutex whose writers run on the loop thread, and sync.RWMutex queues a new reader behind a waiting writer. Re-derived at 371c76a:

Locks taken under pausedMu. The callbacks are the engines' PauseRecv/ResumeRecv closures (engine/epoll/loop.go, engine/iouring/worker.go). They take two locks, one after the other and never nested: detachQMu, then, only when their append took the detach queue from empty to non-empty and after detachQMu is released, WakeFD.mu for reading. So the edges out of pausedMu are pausedMu → detachQMu and pausedMu → WakeFD.mu (R). SetPauser now also takes pausedMu (previous section), but it takes nothing under it.

Neither lock ever waits on anything that could lead back, so no cycle can pass through pausedMu, whatever a caller holds when it takes it:

lock every critical section waits on
detachQMu (both engines) 22 sections: 11 in epoll loop.go, 10 in io_uring worker.go, 1 in transplant_source.go; each is a queue append or slice swap plus an atomic store/swap; Signal is always called after Unlock nothing
WakeFD.mu, read side (Signal) one write(2) on a descriptor New/Set force to O_NONBLOCK nothing (EAGAIN at worst)
WakeFD.mu, write side (Set, Close) fcntl, atomic stores, close(2) only the readers already inside their write(2)

The writers run only on the loop's own thread: Set in run() at start (epoll loop.go:391, io_uring worker.go:992) and when epoll creates its eventfd lazily (loop.go:1807 in OnDetach, :2415 in drainDetachQueue); Close at shutdown (loop.go:3076, worker.go:5413) and in releaseFailedInit (worker.go:5455). None of them is reachable from a pause/resume callback, and the loop thread holds pausedMu only inside requestPause, whose body calls nothing but the callback, and inside SetPauser when a sync-mode upgrade runs on it, which calls nothing at all. So a writer of the RWMutex never waits on pausedMu, holding the write lock or queued for it. A callback's Signal can wait behind a Close, but that Close waits only for the Signals already in write(2).

pausedMu itself is an unexported field of an unexported type, acquired only in engineread.go (requestPause, resumeIfDrained, SetPauser), and the engine packages do not import middleware/websocket.

The forced-interleaving test (TestChanReaderWakeFDWritersNeverWaitOnPausedMu, Linux, in CI) uses the production WakeFD and callbacks built like the engines' own, parks a callback while it holds pausedMu between detachQMu's release and Signal, and forces the three ways a writer and that callback can meet:

  • Close while a pause holds pausedMu: Close must complete with pausedMu still held, and the Signal after it is a no-op.
  • Set while a resume holds pausedMu (epoll's lazy eventfd): Set must complete with pausedMu still held, and the Signal after it lands on the new descriptor.
  • The case the RWMutex adds: another producer is inside Signal's write(2) holding the read lock, Close is queued for the write lock behind it, and only then does the pause callback, holding pausedMu, call Signal, which sync.RWMutex queues behind the waiting Close. The goroutine dump confirms each state before the next step. Every lock in the chain is then held or awaited at once, and it must drain once the in-flight write completes. To hold a reader in write(2) for as long as the test needs, the WakeFD gets a full pipe whose O_NONBLOCK the test clears after New set it, which is a longer read hold than production can produce.

A watchdog turns a deadlock into a failure. Control: a mutant in which both writers wait on pausedMu while holding the write lock (a hook right after w.mu.Lock() in Set and Close) fails all three cases on the watchdog instead of hanging the binary (round3/05-lockorder-control/).

Failing-first evidence

Each test lands one commit before its fix. go test -race -count=1 -v -run '^TestChanReader', strict tally of --- PASS/FAIL/SKIP lines (subtests included). Round 2's rows ran natively; round 3's rows ran on Linux (Docker linux/arm64, golang:1.27, memlock 8 MiB), because the lock-order test is Linux-only (round3/03-failing-first.txt, scripts/failing-first3.sh):

commit what it adds PASS FAIL failing tests
7911b9b #672 test (on the #667 fix) 16 5 TestChanReaderStalePauseConverges + all 4 capacities, each with celeris#672: engine recv paused with nothing buffered …
d94aa53 #672 fix 21 0 —
8bf80f0 panic test 21 4 TestChanReaderCallbackPanicReleasesPausedMu + all 3 callback sites, each with pausedMu is still held after the callback panicked …
a14ebbe deferred unlocks 25 0 —
d3472e7 round-2 head 25 0 —
9122766 SetPauser test (round 3) 25 1 TestChanReaderSetPauserOrderedWithAppend: 2 data races, r.pause (engineread.go:252) and r.resume (:283) against SetPauser (:140/:141)
0a3492d SetPauser fix 26 0 —
651d895 lock-order test 30 0 —
69bcd1c, 371c76a header fix, CI 30 0 —
head's tests on d3472e7's engineread.go — 29 1 SetPauserOrderedWithAppend (2 races)
head's tests on main's engineread.go — 16 14 the #667, #672 and panic tests as in round 2; SetPauserOrderedWithAppend (1 race: main's r.pause read, :222); and the lock-order test, whose harness precondition fails (pausedMu is free with the … callback parked), since main applies the callbacks outside pausedMu and the question it asks does not exist there

TestChanReaderStalePauseConverges lays the #672 interleaving out in one goroutine, in the order the race produces it, so the scheduler decides nothing: the crossing append's channel send and depth snapshot (the two statements Append runs before requestPause), the handler draining to empty through the real Read, then that append's requestPause. Only the gap is manufactured; every statement on either side of it is the production path. The oracle is the quiesced state: nothing buffered, so the engine must not be left paused; the reader's view must match the engine's; and the next delivered chunk must reach the next Read. Anti-vacuity: exactly one pause callback, an empty buffer and zero resumes before the stale pause. It sweeps capacity 1 (lowWater 0), 8, 16 and 256.

TestChanReaderCallbackPanicReleasesPausedMu panics once inside each callback site (the pause in requestPause, the resume in Read, the #672 re-check's resume), recovers as a caller would, requires that it recovered its own panic, and probes the lock with TryLock, so a regression fails instead of hanging.

Mutants

Eleven single-site mutants of the head's engineread.go, each run through go test -race -count=1 -v -run '^TestChanReader' (native, like round 2's; round3/04-mutants.txt, scripts/mutants3.py 371c76a). ABSENT counts the head's result lines that a mutant's run does not print:

mutant killed by ABSENT
M1 delete the #672 re-check StalePause, CallbackPanic 0
M2 re-check resumes but leaves pausedState true StalePause (engine/reader disagreement) 0
M3 re-check clears pausedState without resuming StalePause, CallbackPanic 0
M4 re-check uses < instead of <= StalePause/cap1 (lowWater 0) 0
M5 re-check ignores the spill SURVIVES, because of celeris#705 (below) 0
M6 plain Unlock in requestPause CallbackPanic 0
M7 plain Unlock in resumeIfDrained CallbackPanic 0
M8 pause applied outside the lock (#667) LostResumeConverges, then the binary aborts 5
M9 resume applied outside the lock (#667) ParkedResumeConverges, then the binary aborts 5
M10 SetPauser writes without pausedMu (new) SetPauserOrderedWithAppend (2 data races) 0
M11 requestPause's nil check back above the Lock (new) SetPauserOrderedWithAppend (1 data race) 0

M8 and M9 abort the test binary (review point; the previous table did not say so). Each unlocks and relocks pausedMu around the callback inside a critical section whose unlock is deferred. When TestChanReaderCallbackPanicReleasesPausedMu makes that callback panic, the relock is skipped and the deferred Unlock runs on an unlocked mutex: fatal error: sync: unlock of unlocked mutex, which no recover can catch (M8 in pause-in-requestPause, M9 in resume-in-Read). Every test after that point never ran: 5 of the 26 result lines the head prints in this native run are absent (the panic test, its three subtests, and SetPauserOrderedWithAppend). On d3472e7 it was 4 of 25. Both kills stand, because LostResumeConverges and ParkedResumeConverges failed before the abort. But in these two runs nothing after the abort was tested.

M5 survives because of celeris#705, which this PR does not fix. No test reaches the re-check with chunks spilled and the channel at or below lowWater. The only interleaving I found that does is #705's stranded spill: a chunk spills after the handler has drained the whole channel, and it is never promoted, because only a successful dequeue promotes spill. The !hasSpill() guard cannot fix that in either direction. #705 predates this PR: its single-goroutine repro leaves depth=0 spill=1 with the engine paused, on main 9f4d89b exactly as on this head. The previous revision said it was "reported separately"; it had not been filed. It is now celeris#705 (v1.6.0), to be fixed on top of this PR with that repro as its failing-first test, which is what will kill M5.

CI gate

The websocket step's exact-PASS interlock now covers six tests (the two #667 orderings, the #672 stale pause, the panic release, the SetPauser ordering and the WakeFD lock order) and the ten subtests of the stale pause, the panic release and the lock order. A rename, a skip, or a capacity, callback site or interleaving dropped from a sweep turns the step red. The PASS patterns end at (, never at the name. TestChanReaderSetPauserOrderedWithAppend skips without -race, and the step runs -race.

Proved in the CI shape on 371c76a (round3/06-gate-proof.txt, scripts/prove-gate3.sh): the step body is extracted verbatim from ci.yml and run the way GitHub runs shell: bash (bash --noprofile --norc -eo pipefail), with golang:1.27, -race, memlock 8 MiB and the step's WS484_* env. One container per arm:

arm step rc tally
head 371c76a 0 want 6, passed 6; subtests want 10, passed 10 (37 PASS / 0 FAIL / 0 SKIP)
d3472e7's engineread.go (SetPauser unordered) 1 go test fails: SetPauserOrderedWithAppend (race detected)
the six tests renamed 1 go test exits 0; the interlock fails it (passed 0)
the six tests t.Skip 1 6 SKIP; the interlock fails it (passed 0)
one lock-order interleaving dropped (t.Skip) 1 passed 6 but subtests 9/10; the interlock fails it

Round 2's seven arms (main's engineread.go, the #667 fix alone, arm C without the deferred unlocks, and the rename/skip/sweep arms against the four-test interlock) are in round 2's 06-gate-proof.txt.

TestMeasureWedgeRate is outside ^TestChanReader, so the step never sweeps the harness in; without WS667_WEDGE it skips.

End-to-end A/B on real sockets (review MAJOR 2)

celeris#667 makes an end-to-end A/B a merge precondition, and the previous revision's own real-socket A/B found the #667 half raising round failures. So the shipped head is now measured end to end. The design was pre-registered before the seed was drawn and before any analysed round ran: round3/ab/10-PREREGISTRATION.md, sha256 53fe5e80…, hashed 00:42:07Z. mkschedule-ab3.py refuses to draw the seed unless the file still matches that hash. One round is one container and gives one observation per workload.

  • Arms. Every arm is the head tree 371c76a with only engineread.go replaced:

    The only non-test file main and the head differ in is engineread.go, so this compares main's production code with the head's under identical tests and rig. The head arm is 371c76a rather than d3472e7 because this revision changes engineread.go (the SetPauser ordering). The point is to measure the head that ships.

  • Shapes. The CI shape is memlock 8 MiB, --cpus 4, GOMAXPROCS=4, where io_uring gets one worker. The second shape is memlock 128 MiB, where it gets four. Both run Docker linux/arm64 with golang:1.27.

  • W1 is the backpressure and echo tests in CI's step shape (-race, CI's WS484_* env): TestBackpressurePauseDoesNotCancelInflightSend (96 connections per engine), TestBackpressureInboundSequenceIntegrity (16 per engine variant), TestNativeEngineEcho and TestNativeEngineBackpressure.

  • W2 is the earlier A/B's rig, unchanged: BenchmarkWSEngineBackpressureEcho, -count=3 -benchtime=200x, both engines.

  • n and order. 24 rounds per arm per shape (48 per arm pooled), 192 rounds in all. They run in 24 counterbalanced blocks of 8, each a seeded permutation of the four arms followed by its mirror, with blocks alternating shapes. Power (one-sided Fisher at α 0.05, 48 v 48) is 0.76 against the earlier A/B's middleware/websocket: chanReader applies pause/resume to the engine outside the lock that decides them, so a resume can be lost and the connection stops delivering (found by reading, not reproduced) #667-only effect (1.7% → 16.7% of rounds).

  • Decision rule, per workload. Compare this PR with main on rounds failed, pooled over shapes, with a one-sided Fisher test. BLOCK if p < 0.05. VOID if the A/A moves (two-sided p < 0.05).

workload shape main 9f4d89b A/A (main again) #667 only 3b6c500 this PR 371c76a
W1 tests (-race) m8 0/24 (0.0%) 1/24 (4.2%) 1/24 (4.2%) 0/24 (0.0%)
m128 0/24 (0.0%) 0/24 (0.0%) 0/24 (0.0%) 0/24 (0.0%)
pooled 0/48 (0.0%) 1/48 (2.1%) 1/48 (2.1%) 0/48 (0.0%)
W2 rig m8 2/24 (8.3%) 1/24 (4.2%) 4/24 (16.7%) 1/24 (4.2%)
m128 0/24 (0.0%) 0/24 (0.0%) 5/24 (20.8%) 0/24 (0.0%)
pooled 2/48 (4.2%) 1/48 (2.1%) 9/48 (18.8%) 1/48 (2.1%)
contrast (rounds failed, Fisher exact) W1 tests W2 rig
head vs main (primary, one-sided: head fails more) 0/48 vs 0/48, p=1 1/48 vs 2/48, p=0.879
A/A floor: A/A vs main (two-sided) 1/48 vs 0/48, p=1 1/48 vs 2/48, p=1
#667 only vs main (one-sided; the earlier A/B's contrast) 1/48 vs 0/48, p=0.5 9/48 vs 2/48, p=0.025
head vs #667 only (two-sided) 0/48 vs 1/48, p=1 1/48 vs 9/48, p=0.015
head vs main + A/A pooled (one-sided) 0/48 vs 1/96, p=1 1/48 vs 3/96, p=0.807
head vs main at m8 only (one-sided, secondary) 0/24 vs 0/24, p=1 1/24 vs 2/24, p=0.883
head vs main at m128 only (one-sided, secondary) 0/24 vs 0/24, p=1 0/24 vs 0/24, p=1

Verdicts (pre-registered rules): W1: PASS (no increase in round failures that this n resolves); W2: PASS (no increase in round failures that this n resolves).

W1 tests that failed, rounds per arm (pooled over shapes):

W2 failures per arm (pooled): rounds failed by engine, failed reps, and the failed rounds by kind (POST HOC, 32-SENSITIVITY.txt: start = server not ready within timeout, stall = client N: … i/o timeout):

Per-connection failure classes (W1), every class with a non-zero count: per-round median [min-max], total over connections, two-sided Mann-Whitney U over rounds vs main
test / engine / shape class main 9f4d89b A/A (main again) #667 only 3b6c500 this PR 371c76a p head p #667 only p A/A
pause_cancel / epoll / m8 clientCloseFail 0 [0-0] 0/2304 0 [0-0] 0/2304 0 [0-1] 1/2304 0 [0-0] 0/2304 1 0.317 1
pause_cancel / epoll / m8 closedOK 96 [96-96] 2304/2304 96 [96-96] 2304/2304 96 [95-96] 2303/2304 96 [96-96] 2304/2304 1 0.317 1
pause_cancel / io_uring / m8 clientCloseFail 0 [0-0] 0/2304 0 [0-1] 3/2304 0 [0-2] 3/2304 0 [0-17] 19/2304 0.077 0.153 0.077
pause_cancel / io_uring / m8 closedOK 96 [96-96] 2304/2304 96 [95-96] 2301/2304 96 [94-96] 2301/2304 96 [79-96] 2285/2304 0.077 0.153 0.077
pause_cancel / epoll / m128 clientCloseFail 0 [0-0] 0/2304 0 [0-0] 0/2304 0 [0-0] 0/2304 0 [0-1] 1/2304 0.317 1 1
pause_cancel / epoll / m128 closedOK 96 [96-96] 2304/2304 96 [96-96] 2304/2304 96 [96-96] 2304/2304 96 [95-96] 2303/2304 0.317 1 1
pause_cancel / io_uring / m128 clientCloseFail 0 [0-0] 0/2304 0 [0-0] 0/2304 0 [0-1] 1/2304 0 [0-0] 0/2304 1 0.317 1
pause_cancel / io_uring / m128 closedOK 96 [96-96] 2304/2304 96 [96-96] 2304/2304 96 [95-96] 2303/2304 96 [96-96] 2304/2304 1 0.317 1
inbound / epoll / m8 closedOK 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 1 1 1
inbound / io_uring / m8 closedOK 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 1 1 1
inbound / io_uring/multishot_recv / m8 closedOK 16 [16-16] 384/384 16 [16-16] 368/368 16 [16-16] 368/368 16 [16-16] 384/384 1 1 1
inbound / epoll / m128 closedOK 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 1 1 1
inbound / io_uring / m128 closedOK 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 1 1 1
inbound / io_uring/multishot_recv / m128 closedOK 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 16 [16-16] 384/384 1 1 1

Class totals over all rounds, engines and shapes (connections): main 9f4d89b: pause_cancel 9216 conns, clientCloseFail 0; inbound 2304 conns, no failure class non-zero; A/A (main again): pause_cancel 9216 conns, clientCloseFail 3; inbound 2288 conns, no failure class non-zero; #667 only 3b6c500: pause_cancel 9216 conns, clientCloseFail 5; inbound 2288 conns, no failure class non-zero; this PR 371c76a: pause_cancel 9216 conns, clientCloseFail 20; inbound 2304 conns, no failure class non-zero.

Rounds that started with another lane's container running: 177 of 192. First/last round start: ['2026-09-27T01:06:10Z', '2026-09-27T05:04:18Z'].
Rounds per arm and shape, and how many started with another lane's container running: base/m8 24 (23); base/m128 24 (22); aa/m8 24 (23); aa/m128 24 (22); fix/m8 24 (22); fix/m128 24 (21); head/m8 24 (22); head/m128 24 (22).

Unusable rounds (kept, listed, re-run once): ab3-105-m8-head.log: W2 ran no io_uring (engine excluded?); ab3-128-m8-base.log: W2 ran no io_uring (engine excluded?); ab3-156-m8-fix.log: W2 ran no io_uring (engine excluded?). In all three the io_uring sub-benchmark had failed to start before it could log its workers, so these are start failures that the rule excluded; see the sensitivity view above.

Pre-registered verdicts: W1 PASS, W2 PASS. This PR fails no more rounds than main: W1 0/48 vs 0/48, W2 1/48 vs 2/48 (one-sided p = 0.88). The A/A floor is quiet in both (p = 1). A PASS means no increase this n can resolve. At main's observed 2-4%, power reaches 80% only for a head rate of roughly 18-24% (00-power.txt).

The earlier A/B's finding replicates, and this PR removes it. The #667 fix alone fails 9/48 W2 rounds against main's 2/48 (one-sided p = 0.025; the earlier A/B found 20/120 vs 2/120). This PR fails 1/48 (two-sided p = 0.015 against #667 alone).

What the failures are (POST HOC, 32-SENSITIVITY.txt). Read from each failure's own line, every failed round is one of two kinds:

A deviation the analysis exposed. The usability rule "W2 logged workers= for both engines" marked 3 rounds unusable (105 this PR, 128 main, 156 #667 alone; all at 8 MiB), and they were re-run once, as pre-registered. In all three the io_uring sub-benchmark had failed to start before it logged its worker count, so the rule excluded real failures. Round 105 also carried a W1 start failure. Counting the three originals instead of their re-runs (POST HOC) gives:

No conclusion changes, and the stall counts are the same in both views.

Per-connection failure classes.

Conditions.

  • 177 of the 192 rounds started while another lane's container was running. The host is shared, and every arm is affected about equally (21-23 of 24 per arm and shape).
  • The rounds ran from 01:06Z to 05:04Z. The driver was restarted three times between rounds: twice to hold its slot, because lanes that release and re-acquire back to back had starved it, and once after a process-table exhaustion caused by another lane.
  • Rounds 30 and 73-79 never started a container, so they were re-run in position under the pre-registered interrupted-round rule (round3/MANIFEST.txt).
  • While restarting, I saw the outcome lines of two rounds (28 and 29, both this PR). No other interim outcome was looked at.

Throughput (review MAJOR 2, round 4): no loss above the floor

Why the gate changed, and when. The pre-registered load gate (load1 < 1.5) could not open on this host. Over 31 minutes with no container of any lane running, load1 read 4.3–8.8, from dasd and the maintainer's interactive apps, which no lane may touch (round3/ab/timing/33-LOADGATE.txt).

As the review allowed, the gate was amended in a hashed addendum BEFORE any timing round: round4/11-ADDENDUM.md, sha256 fac26e72…, hashed 06:17:06Z. The first timing round started at 08:45:23Z.

  • What replaced the gate:

    • one hold of the laptop timing lock;
    • no container and no slot of any lane at each round's start;
    • each round logs its start load1 and the top CPU processes.

    Noise is left to the design, as pre-registered: counterbalanced mirrored blocks, and each cell's own A/A floor.

  • What did not change: the arms and their trees, both shapes, the W2 command (byte for byte), the 192-round seeded order, the metric, the BLOCK rule and the analysis script (ab3stats.py --timing, unchanged, sha256 be9bf7f4…).

  • The declared secondary (the review's suggestion): after W2, each round's container runs W2X. W2X is the same rig with every op timed, and a second pass with the mutex profile on. It uses a copy of the arm's tree plus one test file, byte-identical in every arm and not part of this PR.

The run. 192 rounds, 08:45–10:02Z, no unusable round. At every round's start no other container was running and no slot was held.

  • One whole-block pause: 09:03:53–09:15:24Z, between blocks 7 and 8. The orchestrator asked for the timing lock for fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674, and the addendum allowed a pause only at a block boundary.

  • Load: start load1 was 2.3–5.4 (median 3.65). It was balanced across arms: against main, Mann-Whitney p ≥ 0.37 in every arm and shape.

  • Foreign load, reported after the run. Another lane's native go test -race -p 2 runs overlapped this hold twice, because its host check did not gate on the timing lock (round4/FOREIGN-LOAD-NOTE-r4h2-20260927T0914Z.txt):

    • 09:14:54–09:20:01Z: timing rounds 57–66, blocks 8–9;
    • 10:06:12–10:06:36Z: replication rounds 1–5, block 1.

    The note asks for those rounds to be voided and re-taken. The pre-registration has no rule for foreign native load, so the pre-registered analysis keeps them. POST HOC, dropping those whole blocks changes no verdict (38-FOREIGN-LOAD-SENSITIVITY.txt):

    • the matrix still PASSES in all four cells (io_uring at 128 MiB: +2.20%, p = 0.12);
    • the replication reads +0.26% (p = 0.15);
    • the pooled io_uring view at 128 MiB becomes +0.60% [-0.41, +2.28] (p = 0.06).

Pre-registered verdicts: all four cells PASS (round4/31-THROUGHPUT.txt). One round is the median of 3 reps, with 24 rounds per arm and cell. A cell BLOCKs if this PR is slower than main by more than that cell's A/A floor at one-sided p < 0.05.

engine memlock main ns/op (median of rounds) this PR vs main [95% round bootstrap] A/A floor one-sided p (this PR slower) verdict #667 alone vs main
epoll 8 MiB 894,262 +1.58% [-1.70, +3.95] 1.90% 0.32 PASS +2.80% (p=0.0079, 4 failed rounds)
io_uring 8 MiB 786,812 -2.26% [-5.86, +3.25] 3.08% 0.9 PASS -2.16% (p=0.84, 2 failed rounds)
epoll 128 MiB 903,878 +0.58% [-1.65, +3.12] 0.13% 0.2 PASS +1.76% (p=0.035, 3 failed rounds)
io_uring 128 MiB 909,026 +2.29% [-0.34, +6.55] 0.03% 0.059 PASS +2.50% (p=0.097, 2 failed rounds)

The #667-only arm's failed rounds are stalls, which the rank test counts as slowest.

MDE (the smallest slowdown of this PR the one-sided test detects with 80% power, from the base and A/A spread; reported afterwards, not a gate): epoll 8 MiB +3.5%; io_uring 8 MiB +7.5%; epoll 128 MiB +3.0%; io_uring 128 MiB +3.0%.

The borderline cell was replicated. io_uring at 128 MiB (four workers) read +2.29% at p = 0.059. The declared load sensitivity, which drops rounds above the 90th percentile of start load, read +3.37% at p = 0.019 in that cell. The primary is not a BLOCK. The sensitivity view would meet the rule, but it was declared with no decision. And one run discriminates nothing.

So I pre-registered a replication of that cell alone, before its seed: round4/12-CONFIRM.md, sha256 2a77764f…, hashed 10:06:05Z.

  • Design: main, A/A and this PR at 128 MiB, 64 rounds per arm (192 rounds), with W2 only, the same rule and the same script.
  • Power, computed after the run (round5/53-MDE-CONFIRM.txt; no verdict depends on it). ab4-secondary.py's MDE function gives +2.0% at 64 rounds per arm, the figure 12-CONFIRM.md quotes. With 4,000 draws per point, the pre-registered rule reaches 80% power between +1.75% and +2.0%. At round 4's +2.29% its power is 0.93, not the "about 80%" that 12-CONFIRM.md states. That sentence had no saved computation behind it, and it was wrong. The hashed file is left as it was.
  • Result: it ran 10:06–10:13Z with no unusable round (round4/35-CONFIRM.txt).
cell (replication, 64 rounds per arm) this PR vs main [95% round bootstrap] A/A floor one-sided p verdict
io_uring, 128 MiB (primary) +0.40% [-0.81, +2.29] 0.97% 0.096 PASS
epoll, 128 MiB (secondary) -1.10% [-2.18, -0.16] 0.72% 0.99 PASS

So there is no loss above the floor. The one open question is a small cost on io_uring with four workers. Both runs of that cell pooled, 88 rounds per arm, give +1.40% [-0.20, +2.54], one-sided p = 0.024, against an A/A floor of 0.56% (round4/36-CONFIRM-POOLED.txt). That view was declared secondary, and it is biased upward, because round 4 is what flagged the cell. The unbiased test is the replication alone, and it passes. Without the blocks that overlapped foreign load, the pooled view is +0.60% [-0.41, +2.28] (post hoc). Either way, a cost of up to about 2.5% on io_uring with four workers is not excluded. At 8 MiB (one worker, the CI shape) this PR read -2.26%.

Worker stall: the mechanism the review names is real, and small on average (declared secondary; round4/34-SECONDARY.txt, with the post hoc detail in 37-WORKERSTALL-DETAIL.txt). BenchmarkAB4EchoMutex counts only waits that parked, attributed to the lock holder's stack. With the mutex profile on, the engine worker in requestPause parks on a pausedMu that the handler holds across the resume callback:

engine, memlock main: stalls / 14,400 ops, mean wait A/A #667 alone this PR per-round median, this PR vs main (one-sided p)
epoll, 8 MiB 3, 4 µs 4, 5 µs 7 / 10,200, 114 µs 22, 105 µs 18 vs 0 ns/op (p = 1.6e-4)
io_uring, 8 MiB 2, 239 µs 3, 1 µs 29 / 13,800, 102 µs 33, 82 µs 192 vs 0 ns/op (p = 7.5e-7)
epoll, 128 MiB 2, 21 µs 3, 2 µs 3 / 7,200, 9 µs 23, 89 µs 30 vs 0 ns/op (p = 3.5e-5)
io_uring, 128 MiB 3, 3 µs 4, 3 µs 19 / 10,800, 143 µs 30, 180 µs 261 vs 0 ns/op (p = 3.6e-7)

The #667-only arm has fewer ops because its stalled reps failed.

  • How often and how long: about once per 440–650 ops, a wait of 80–180 µs on average (the largest mean over one rep is 0.47 ms). On main the event is roughly ten times rarer, and in three of the four cells it lasts a few µs. Every connection on that worker waits that long.
  • What it costs: amortised, 140–375 ns per op, where an op takes 0.8–0.9 ms. That is 0.02–0.04%, far below anything the throughput test can see.

The cost comes from holding pausedMu across the callbacks, which is what #667 requires. Making those waits rarer would mean changing how the callbacks are ordered, and that is not this PR.

Tail: not resolved at this n, and a small tail cost in epoll is possible. These are all of the A3 declared tail metrics (round4/34-SECONDARY.txt); the post hoc detail is in round5/52-TAIL.txt, from scripts/tail-detail.py. Order: main / A/A / #667 alone / this PR. Per-round contrasts are this PR vs main (one-sided Mann-Whitney, uncorrected).

engine, memlock pooled p99.9 op time (µs) ops ≥ 10 ms, pooled (one-sided binomial) per-round p99 per-round max per-round slow ops (> 4x the round's median) / ops ≥ 10 ms
epoll, 8 MiB 3,058 / 2,583 / 10,618 / 14,196 13 / 10 / 15 / 16 (p = 0.36) +1.1% (p = 0.23) +28% (p = 0.59) p = 0.39 / 0.31
io_uring, 8 MiB 4,283 / 3,702 / 3,631 / 4,135 0 / 0 / 0 / 0 +1.1% (p = 0.59) −2.2% (p = 0.72) p = 0.46 / 1
epoll, 128 MiB 1,814 / 1,888 / 4,097 / 12,199 10 / 9 / 9 / 17 (p = 0.12) +3.2% (p = 0.19) +345% (p = 0.042) p = 0.095 / 0.095
io_uring, 128 MiB 1,735 / 1,685 / 1,871 / 1,926 6 / 4 / 7 / 5 (p = 0.73) +1.5% (p = 0.21) +7.6% (p = 0.37) p = 0.63 / 0.63

Each arm pools 600 ops from each of 24 rounds, 14,400 ops (the #667-only arm has 13,200 in the two cells where two of its rounds lost a rep).

  • Why p99.9 jumps. At 14,400 ops, p99.9 is the 15th slowest op, so it mostly records whether 15 ops passed 10 ms.

    • In epoll at 128 MiB this PR has 17 ops at or above 10 ms, so its 15th slowest is 12.2 ms. Main has 10, so its 15th slowest is 1.8 ms.
    • In epoll at 8 MiB main has 13, and its 14th slowest op is 9.9 ms. One more op over 10 ms would move main's p99.9 from 3.1 ms to 9.9 ms.
    • A round bootstrap of the pooled p99.9 runs from 1.5–1.9 ms to 13–17 ms in every epoll arm, main's and the A/A's included.

    So the pooled p99.9 moves by about 10 ms on a difference of a few ops. The counts are the steadier reading of the same data.

  • epoll at 128 MiB: every tail metric leans against this PR, and one declared contrast reaches p < 0.05.

  • epoll at 8 MiB is the same epoll setup. Every round of both shapes logs epoll: workers=4, and memlock limits only io_uring. Here the declared counts lean the same way, but less: 16 vs 13 ops at or above 10 ms (p = 0.36). The 2 ms count shows nothing: 19 vs 18.

  • io_uring shows no lean at 2, 4 or 10 ms in either shape. For example, ops at or above 2 ms are 32 vs 33 at 8 MiB and 13 vs 10 at 128 MiB (p ≥ 0.34).

  • At 1 ms, about the 80th percentile of an op, the count describes the body of the distribution, not its tail. There the pooled binomial is not usable: in epoll at 8 MiB the A/A arm has 2,411 such ops against main's 2,898, a 17% gap between identical code. Per round, this PR vs main gives p ≥ 0.21 in every cell.

  • So the tail is not resolved at this n. A small tail cost in epoll is possible, most visible at 128 MiB; nothing here points to one on io_uring. If the cost is real, its likely source is the hold of pausedMu across the callbacks that the worker-stall numbers above measure.

Stalls, again (descriptive; the stall class is post hoc). A W2 round stalls when a client's 60 s read deadline expires, which is the #672 wedge. Every count and p-value here is in round5/51-STALLS.txt, written by scripts/stall-counts.py from the raw round logs. It reproduces round 3's 32-SENSITIVITY.json counts exactly.

Measurement: celeris#672 wedge rate on the rebased head (round 2)

Measured on d3472e7's engineread.go. 371c76a's differs from it only in the SetPauser ordering, and the harness calls SetPauser once, before it starts its goroutines, so these numbers stand for the head. They are a unit-level harness; the real-socket measurement is the section above.

The same harness and round shape as the previous revision, now committed. The escape hatch is removed, so one wedge ends one connection. A wedge is detected structurally: the engine is paused, the buffer is empty, and the reader makes no progress across 512 runtime.Gosched() yields. No wall clock is involved. Each round is 16 connections, one container and one process, and counts as ONE observation. Setup: golang:1.27, --cpus 4, GOMAXPROCS=4, memlock 8 MiB, no -race, Docker linux/arm64.

Pre-registered before the seed was drawn and before any round ran:

capacity arm wedges/kiloEdge (round median) connections wedged [round-cluster 95%] rounds with a wedge median ops to first wedge edges/op
cap256 (default) base (main) 0.1901 560/640 (87.5%) [84.7, 90.2] 40/40 206,252 0.013119
aa (A/A) 0.2043 579/640 (90.5%) [88.0, 93.0] 40/40 225,746 0.013064
#667 alone 0.6174 623/640 (97.3%) [95.8, 98.6] 40/40 54,829 0.012896
this PR 0 0/640 0/40 — 0.013040
cap16 base (main) 0.0436 515/640 (80.5%) [76.4, 84.5] 40/40 62,304 0.189167
aa (A/A) 0.0453 498/640 (77.8%) [73.9, 81.7] 40/40 43,734 0.189827
#667 alone 0.6096 640/640 (100%) 40/40 2,271 0.186068
this PR 0 0/640 0/40 — 0.189559
contrast (round level, n=40v40) cap256 cap16
A/A floor: aa vs base x1.07 [0.84, 1.33], p=0.53 x1.04 [0.79, 1.23], p=0.81
#667 alone vs base x3.25 [2.04, 4.62], p=4.1e-08, Holm reject x14.0 [9.0, 17.5], p=1.7e-10, Holm reject
this PR vs base 0 vs 0.19, U=0, p=1.9e-16, Holm reject 0 vs 0.044, U=0, p=1.9e-16, Holm reject
this PR vs #667 alone 0 vs 0.62, p=1.9e-16 0 vs 0.61, p=1.9e-16
edge gate (armc / fix / aa vs base) -0.61% / -1.70% / -0.42%: all pass +0.21% / -1.64% / +0.35%: all pass

Across 80 rounds this PR wedged 0 of 1,280 connections, over 32.1 million watermark edges. By the pre-registered verdicts:

Controls:

  • The A/A floor is quiet at both capacities, so the contrasts are not the apparatus measuring itself.
  • The detector can report zero, and it did. The conservative fallback never fired. Arm C's connections inspected up to 673 transient paused-and-empty states each, against a fallback at 4,096, and every one resolved.
  • No order effect at p<0.05. The pooled slot contrast is +5.9% (p=0.98) at cap256 and -10.0% (p=0.69) at cap16. The largest within-arm slot contrast is base at cap16, -24.7% (p=0.058).
  • Anti-vacuity. Every connection must record edges>0 && pauses>0 or the round fails, and all 320 rounds were usable: GO_RC 0, one PASS line, 16 connection lines.

Two things I have to disclose.

  1. I looked at an interim analysis mid-matrix. At 90 of 160 cap256 rounds (n=23 per arm), the matrix stalled behind another lane's hold on the shared Docker slot, and I ran the analysis to check the pipeline. It showed the edge gate for arm C vs base at -11.8%, which is VOID. I wrote down what that did and did not change before the matrix finished (deviation D1): nothing about any verdict, a post-hoc per-connection edge-rate comparison to be shown next to the gate, and a separate trigger-match check (F2) to be run whatever the final gate read. At n=40 the pre-registered gate passes at -0.61%. The interim gap came from the first 23 arm C rounds (edges/op 0.01154 against 0.01315 for the last 17); base barely moved over the same split. Per connection, arm C vs base differs by +0.17% at cap256 and +1.53% at cap16 (post-hoc).
  2. F2 does not confirm a trigger match at cap256, and does not refute one either. F2 is BenchmarkChanReaderContended, whose escape hatch runs every arm for the same op count, at n=10 per arm. There arm C is -6.64% vs base, but the benchmark's own A/A moved -4.78% at that n. cap16 is confounded by base degeneration (8/10 rounds), as fixed in advance. This changes no verdict: the per-edge metric already normalises for edge rate, and a 7% edge-rate difference cannot turn 560/640 wedged connections into 0/640.

Whole suites, main vs this branch

go test -race -count=1 -v, full packages, one container per (arm, suite), golang:1.27, --cpus 4, seccomp unconfined, base and branch alternating which runs first. Strict tally: only --- PASS/FAIL/SKIP: lines, subtests included; a SKIP is counted as absent, never as a PASS.

suite main 9f4d89b: rc, PASS / FAIL / SKIP this branch d3472e7: rc, PASS / FAIL / SKIP PASS→FAIL PASS→not-PASS status changes
./middleware/websocket, memlock 8 MiB (CI's WS484_* env) 0, 214 / 0 / 1 (864s) 0, 225 / 0 / 2 (856s) 0 0 11 new chanReader tests/subtests absent→PASS; TestMeasureWedgeRate absent→SKIP
./middleware/websocket, memlock 128 MiB (CI's WS484_* env) 0, 214 / 0 / 1 (883s) 0, 225 / 0 / 2 (882s) 0 0 11 new chanReader tests/subtests absent→PASS; TestMeasureWedgeRate absent→SKIP
./engine/epoll (its tests import middleware/websocket) 0, 99 / 0 / 5 (26s) 0, 99 / 0 / 5 (25s) 0 0 0 new chanReader tests/subtests absent→PASS
./engine/iouring, memlock 8 MiB (one worker) 0, 281 / 0 / 4 (126s) 0, 280 / 0 / 5 (129s) 0 1 0 new chanReader tests/subtests absent→PASS; TestListenRefusesToStartWhenNoWorkerReportsItsAddress PASS→SKIP
./engine/iouring, memlock 128 MiB 0, 285 / 0 / 0 (127s) 0, 285 / 0 / 0 (127s) 0 0 0 new chanReader tests/subtests absent→PASS
./engine/iouring, memlock 8 MiB, second run (branch first) 0, 280 / 0 / 5 (121s) 0, 280 / 0 / 5 (122s) 0 0 0 new chanReader tests/subtests absent→PASS

PASS→FAIL: 0. PASS→SKIP/absent: 1.

  • The one SKIP main has in both websocket runs is TestHubBroadcastFormatsOnce. It skips on the branch too. The branch's extra SKIP is the wedge harness, which skips without WS667_WEDGE.
  • The PASS→SKIP is TestListenRefusesToStartWhenNoWorkerReportsItsAddress in ./engine/iouring at 8 MiB. It skips when io_uring is unavailable, and on the branch run io_uring_setup returned ENOMEM (cannot allocate memory … current=8388608 bytes). The locked memory was momentarily exhausted in that process, and the log shows it. The cause is environmental, not this PR: ./engine/iouring has no diff between the two arms and does not depend on middleware/websocket (go list -deps -test finds 0 matches), so both arms run the same test code. It passed on the branch at 128 MiB. A second full run of both arms at 8 MiB (branch first, iouring-m8-rerun) has it SKIP on the branch and SKIP on main (strict tally). Main skips it too, with the same io_uring_setup ENOMEM, so the skip is not the branch's. Across the four 8 MiB runs it passed 1 time and skipped 3 times, each skip on ENOMEM. At 8 MiB on this host it skipped more often than it ran, whichever arm ran it, and a skip there is a hole in coverage, not a verdict.

Round 3, RULE 21 before the push: the whole ./middleware/websocket suite on 371c76a, the same shape as the first row above (-race -count=1 -v, memlock 8 MiB, CI's WS484_* env, one container): rc 0, 230 PASS / 0 FAIL / 2 SKIP in 885 s (round3/07-suite/). Against round 2's runs of main and of d3472e7 (round3/07-suite-compare.txt): PASS→FAIL 0, PASS→not-PASS 0. The only new result lines are the five new tests and subtests, all PASS. The two SKIPs are the same as before: TestHubBroadcastFormatsOnce (skips on main too) and the wedge harness (skips without WS667_WEDGE).

celeris#633's coverage reproducer, main vs this branch (round 2)

Pre-registered before any round ran (60-cover633/10-PREREGISTRATION.md). The input is celeris#633's cheap reproducer: the CI step go test -race -covermode=atomic -run '^TestBackpressure' ./middleware/websocket/..., with CI's WS484_* env at memlock 8 MiB. On a GitHub runner it failed TestBackpressurePauseDoesNotCancelInflightSend on both engines (4 close-timeouts each). One round is one container and one go test process. There are 10 rounds per arm, in blocks of four: a seeded permutation of the two arms, then its reverse. Every round ran on Docker linux/arm64 with --cpus 4. Pre-registered deviations from the CI command: -v, -timeout=900s, and arm64.

main 9f4d89b this branch d3472e7
rounds in which TestBackpressurePauseDoesNotCancelInflightSend failed (primary) 10/10 10/10
/epoll subtest failed 10/10 10/10
/io_uring subtest failed 8/10 7/10
epoll closeTimeout per round, median [min–max] of 96 conns 0 [0–3] 0 [0–29]
epoll otherWriteErr per round, median [min–max] of 96 conns 43 [9–53] 41 [31–59]
epoll clientCloseFail per round, median [min–max] of 96 conns 88.5 [13–94] 80 [15–94]
epoll closedOK per round, median [min–max] of 96 conns 7.5 [2–80] 16 [2–53]
io_uring closeTimeout per round, median [min–max] of 96 conns 2 [0–58] 2.5 [0–69]
io_uring otherWriteErr per round, median [min–max] of 96 conns 1 [0–7] 0.5 [0–3]
io_uring clientCloseFail per round, median [min–max] of 96 conns 34.5 [2–74] 64 [2–85]
io_uring closedOK per round, median [min–max] of 96 conns 49.5 [6–86] 16.5 [5–94]
TestBackpressureInboundSequenceIntegrity passed 10/10 9/10

Primary: this branch 10/10 vs main 10/10 failed rounds, two-sided Fisher p = 1: no move this n can resolve.
The reproducer fails on main in 10 of 10 rounds on this host, so the failure predates this PR. This PR does not fix it. With main failing every round, the design would have called a move only if this branch had failed 5/10 rounds or fewer (00-power.txt). A null result at this n does not show the rate is unchanged.

TestBackpressureInboundSequenceIntegrity failed once, on this branch (cover633-07-head.log), in io_uring/multishot_recv, with inbound_sequence_linux_test.go:132: server not ready within timeout. That is the test helper's 30 s wait for the server to report a listening address. The io_uring server in that sub-run logged its memlock worker cap and never logged that it was listening, so no connection was upgraded and no frame was sent. It is not an integrity verdict, and it did not run the code this PR changes (chanReader exists only after an upgrade). Main 0/10 vs this branch 1/10, Fisher p=1. The helper discards StartWithListenerAndContext's error, so the log cannot say whether the start failed or hung (now celeris#706; the round-3 A/B saw the same start failure in every arm at 8 MiB).

The 5 engine subtests that PASSED (both arms together) had clientCloseFail from 2 to 85 of 96 connections (median 23). That is the counter celeris#623 reports as printed but never asserted.

Secondary, per engine, two-sided Mann-Whitney U over rounds, this branch vs main: epoll closeTimeout p=0.43; epoll clientCloseFail p=0.73; epoll otherWriteErr p=0.76; io_uring closeTimeout p=0.91; io_uring clientCloseFail p=0.33; io_uring otherWriteErr p=0.19. None is below 0.05.

Deviations and conditions (read from the round logs):

  • Round 14 ran twice. The first attempt (void/cover633-14-base.attempt1.log, started 2026-09-26T17:49:21Z) stops after the first test's listener line. It has no result line and no ROUND_RC, and the driver logged no end for it. The host's once-a-minute process sampler (_locks/proc-sampler.log) read 1038 user processes at 17:39:06Z and then logged nothing until 18:39:23Z. That is consistent with fork failing, as it had twice earlier that afternoon (the lock tool's fork: Resource temporarily unavailable lines in driver.txt). Docker's VM clock in that log reads 7 minutes behind the host. The attempt is void as apparatus, not an observation. It was re-run with the same tree, command and slot discipline, and only the re-run is analysed. Counting the void attempt as one more failed main round instead changes nothing (p=1; not pre-registered).
  • The rounds ran in 6 stretches, not back to back (rounds 1–4 15:59–16:11Z; rounds 5–8 16:42–16:52Z; rounds 9–13 17:30–17:49Z; rounds 14–16 22:14–22:19Z; round 17 22:51–22:52Z; rounds 18–20 23:40–23:44Z, 2026-09-26). The driver log shows why: it waited for the shared Docker slot or behind another lane's timing lock, the lock tool failed to fork during the process-table exhaustions, and the lane's agent stalled before the last rounds.
  • Block 4 (rounds 13–16) straddles a gap: 13 (branch) before it, 14 (main), 15 (main), 16 (branch) after it. Its arms are therefore not balanced in time across that gap.
  • Block 5 (rounds 17–20) straddles a gap: 17 (branch) before it, 18 (main), 19 (main), 20 (branch) after it. Its arms are therefore not balanced in time across that gap.
  • Round 13 took 653 s of wall time for a 74 s go test package run. The delay came after its last result line, and its results are complete, so it counts as pre-registered.
  • Rounds 3, 5, 6, 7, 8, 9, 19 started while another lane's container was running. The slot lock allows two. That applies to both arms: 4 main and 3 branch rounds.
  • Pre-registered, not deviations: -v, -timeout=900s instead of 120 s, and linux/arm64 instead of CI's amd64.

Not in this PR

Reproducing

  • The wedge harness is in the tree: middleware/websocket/engineread_wedge_linux_test.go. Its header gives the command for one observation and builds the base and middleware/websocket: chanReader applies pause/resume to the engine outside the lock that decides them, so a resume can be lost and the connection stops delivering (found by reading, not reproduced) #667-only arms with git show 9f4d89b:… and git show 3b6c500:…. Before this revision the header described the middleware/websocket: chanReader applies pause/resume to the engine outside the lock that decides them, so a resume can be lost and the connection stops delivering (found by reading, not reproduced) #667-only arm as "this tree's engineread.go with the re-check deleted", which is not the arm the A/B measured (review point).
  • Round 3's evidence and scripts are kept with the measurement logs under lane-20260926/round3/. Its MANIFEST.txt maps every output to the script that produced it:
    • one round: ab3-round.sh, which starts the container through dock.sh (rounds 1–27), dock2.sh (28–29) and, under a held slot, dock-held.sh (30 on);
    • the correctness matrix: run-ab3-matrix.sh ran rounds 1–29, run-ab3-matrix2.sh 30–32, run-ab3-matrix3.sh 33–72, and run-ab3-matrix4.sh 73–192 plus the three re-runs. Each restart is logged in ab/driver.txt;
    • the schedule and power: mkschedule-ab3.py draws the seed and the schedules, and refuses to run unless the pre-registration still matches its hash. power-ab3.py computes the power;
    • timing: run-ab3-timing.sh, then run-ab3-timing2.sh, only waited on the load gate. No timing round ran under either. Round 4's matrix ran under run-ab4-timing.sh (below);
    • the analysis:
      • ab3stats.py makes the pre-registered round-failure verdicts and contrasts and the per-connection classes. With --timing it makes round 4's throughput verdicts. ab3-section.py turns its output into the A/B tables;
      • ab3-sensitivity.py (POST HOC) makes the start and stall kinds, round 3's stall p-values and the usability-rule sensitivity (ab/32-SENSITIVITY.txt);
      • loadgate-summary.py makes the load-gate figures (ab/timing/33-LOADGATE.txt);
    • the rest: failing-first3.sh, mutants3.py, lockorder-control.sh, prove-gate3.sh, run-suite3.sh + suite-compare.py, fetch-ci3.sh, and assemble3.py (the body).
  • Round 4's are under lane-20260926/round4/ (MANIFEST.txt):
    • 11-ADDENDUM.md + .sha256, and 12-CONFIRM.md + .sha256 (hashed before the seed);
    • the throughput matrix: run-ab4-timing.sh → ab4-round.sh (with ab4-warmup.sh);
    • the replication: mkschedule-confirm.py → run-ab4-confirm.sh → ab4c-round.sh;
    • the analysis: round 3's ab3stats.py --timing for both primaries, ab4-secondary.py (tail, worker stall, load, MDE), confirm-pooled.py and tp4-tables.py;
    • POST HOC: workerstall-detail.py (the worker-stall table's events and mean waits, 37-WORKERSTALL-DETAIL.txt) and foreign-sensitivity.py (38-FOREIGN-LOAD-SENSITIVITY.txt);
    • the body: assemble4.py;
    • the W2X rig: apparatus/zz_ab4_secondary_linux_test.go.
  • Round 5's are under lane-20260926/round5/ (MANIFEST.txt). They add no measurement and re-read round 3's and round 4's logs:
    • stall-counts.py → 51-STALLS.txt, every stall count and its Fisher test;
    • tail-detail.py → 52-TAIL.txt, the tail in full;
    • mde-confirm.py → 53-MDE-CONFIRM.txt, the replication's power;
    • assemble5.py, the body.
  • Round 2's are unchanged: run-wedge4-all.sh → run-wedge4.sh, mkschedule-wedge4.py, power.py, report-wedge4.py → wedgestats4.py, failing-first.sh, mutants.py, prove-gate.sh + extract-step.py, run-suites.sh + tally.py, mkschedule-cover633.py → run-cover633.sh → stats-cover633.py (with control-cover633.sh), fetch-ci.sh.

Changes

Test Plan

  • Unit tests added/updated: middleware/websocket: a pause decided on a stale depth snapshot wedges the connection permanently when the handler drains to empty first (distinct from #667) #672, panic, SetPauser-ordering and lock-order tests, each failing on the commit before its fix (the lock-order test's control is a mutant, since nothing in the tree ever made it fail)
  • They run in CI, and the interlock is proved in the CI shape (round 3: five arms, including rename, skip and a dropped interleaving)
  • Tested on Linux: Docker, golang:1.27, linux/arm64, memlock 8 MiB and 128 MiB; end to end on real sockets with both engines (the A/B above)
  • Whole ./middleware/websocket suite on 371c76a before the push (RULE 21): 230 PASS / 0 FAIL / 2 SKIP
  • Throughput (round 4), with no code change since 371c76a. The pre-registered 192-round matrix PASSES in all four cells. The pre-registered 192-round replication of io_uring at 128 MiB PASSES.
  • CI on 371c76a (run 36284528042, attempt 1): all 9 jobs green, plus CodeQL. The websocket step's own log reads celeris#667/#672 regression tests: want 6, passed 6; subtests want 10, passed 10. Per-test result lines exist only for go test -v runs, so the strict tally covers those steps only: Unit's websocket and named-test steps 138 PASS / 0 FAIL / 0 SKIP, Conformance 119/0/0, Adaptive 103/0/0, io_uring init-failure 3/0/0. The Unit job's root race step and its four middleware sub-module steps, and the four Driver Conformance steps, run without -v and print no per-test lines, so a SKIP in them cannot be counted from the log. Round 2's "0 SKIP" had the same limit and did not say so (review point).

Tested on: [ ] std [x] epoll [x] io_uring — [x] amd64 (compile + CI) [x] arm64

Release notes

  • Breaking change? (label breaking)
  • Labeled for release notes (bug)

Summary by CodeRabbit

  • Bug Fixes
    • Improved WebSocket connection reliability under backpressure. Pause and resume behavior now stays consistent during concurrent activity, helping prevent connections from becoming stuck or failing to resume reading after buffer pressure clears.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Not merging this alone: measured, it makes #672 worse at the production default

The corrected A/B answered the question it was rebuilt to answer, and the answer blocks the
merge. Recording it here so the decision is on the PR rather than only in a log.

Throughput: inconclusive, and the old headline is withdrawn. At round-level analysis
(one container = one process = one observation, -count reps collapsed to a round median before
any test) the end-to-end shapes are null against their own slot-matched A/A floors:
w2-m8 io_uring +2.67% p=0.456 (floor +0.71%), w2-m128 io_uring +0.11% p=0.539, both epoll
cells inside ±0.1%. The previously published +45.03%/+44.37% cap16 figures are formally
withdrawn: they were selection on the dependent variable (the filter dropped 32/60 and 49/100
base rounds and 0 fix rounds), and the two unfiltered replicates disagree in sign
(-2.73% p=0.759 and +12.32% p=0.0034) while both are confounded by an edge-rate mismatch. A
low-single-digit cap256 cost remains plausible but unestablished; anything larger is excluded.

The wedge rate is the decisive result, and it goes against this PR. Measured directly per
connection, with the benchmark's escape hatch removed so one wedge ends one connection — which is
what production sees, since production has no hatch:

capacity base this PR ratio n
cap16 0.0427 0.5555 wedges/kiloEdge ×13.0 [10.7, 18.7], p=1.4e-14 40v40 rounds
cap256 (production default) 0.1573 0.2560 ×1.63 [1.35, 1.97], p=8.4e-12 40v40 rounds

Connections wedged at cap256: 556/640 → 615/640. Median ops-to-first-wedge: 285,066 → 170,726.
The pre-registered edge-rate gate passes in both rows (−0.17% and −0.51%), so these are
like-for-like. Independent ecological corroboration from the other harness, real sockets:
end-to-end round failures 2/120 (1.7%) on base → 20/120 (16.7%) on the fix, against a same-binary
A/A floor of 4/240 (1.7%); Fisher fix-vs-base p=6.0e-05, fix-vs-floor p=2.6e-07, base-vs-floor
p=1.0.

And the remedy is already measured. Arm C — this fix plus a #672 stale-pause re-check inside
requestPause — wedged 0 of 640 connections at cap16 and 0 of 640 at cap256, at matched edge
rates, p≈1.9e-16 against the fix arm. The detector can report zero, so its non-zeros mean
something.

So #667 and #672 should land together. Fixing the reordering while leaving the stale-pause hole
open is a net regression at the default capacity, even though the reordering fix is correct in
itself.

Also to address before this merges

  • A rep-level exclusion inside the row labelled as the non-excluded view (major): a failed w2
    round still contributes a round median computed over only its surviving reps, while being
    counted in the n=60 "including failed rounds" primary.
  • pausedMu is released without defer across the engine callback, on a path where panics
    are recovered — so a recovered panic converts into a permanently wedged connection. That is the
    very failure mode this PR is arguing about, introduced by its own fix shape.
  • The wedge harness that produced the decisive numbers is untracked, so the result is not
    reproducible from the repository. Commit it.
  • Comment and body cite engine/epoll/loop.go:1672-1703 for the Swap closures; the rebase moved
    them.
  • Wilson intervals pool 640 connections across 40 rounds as if independent — the same clustering
    error as M4, recurring in an interval instead of a p-value.
  • The pre-registered MDE decision rule is absent from the body; record it or record its
    withdrawal as a deviation.

Credit where due: M1–M5 are genuinely resolved, and both re-reviewers re-derived the numbers
independently rather than taking the report's word.

…the lock that decides it

chanReader decides to pause under pausedMu and applies the decision to the
engine after releasing it (engineread.go:230-232), and Read's resume branch
does the same in reverse (engineread.go:311-313). The engine's closures are
Swap-based and ignore the previous value (engine/iouring/worker.go:2200-2231),
so whichever callback reaches the engine last wins outright and nothing
reconciles.

Two failing-first tests, both engine-free and both deterministic:

  TestChanReaderLostResumeConverges  - the appender parks inside pause() after
    unlocking; the handler drains to lowWater, clears pausedState and applies
    resume() into an engine that is not paused yet; the pause then lands. The
    engine's recv ends up PAUSED while the reader believes it is RUNNING, and
    the reader is edge-triggered, so it never resumes that connection again.

  TestChanReaderParkedResumeConverges - the symmetric direction. The handler
    parks inside resume(); a later pause lands first and the resume overwrites
    it. The engine keeps delivering while the reader believes it is paused, so
    backpressure is silently gone.

No timing on the failing path: the order of the two applications is fixed by a
channel close that strictly happens-before the parked callback's Swap, so the
outcome is forced rather than sampled. Both tests carry anti-vacuity guards
(exact pause/resume callback counts, and proof that the park actually
happened), so a run in which the watermark crossing never fired fails as a
harness error instead of passing quietly.

The park is scheduled by a TryLock probe of pausedMu, not asserted on, so the
tests do not mandate one particular fix and cannot deadlock on a build that
closes the window: if a callback is ever invoked with the deciding lock held,
the park is skipped and the same convergence assertion runs.

Measured on 6136554: 40/40 FAIL over 20 iterations of each test under
-race, 0 passes. With r.pause()/r.resume() moved inside pausedMu: 40/40 PASS.
The fix itself is deliberately NOT included here - issue #667 requires the
contention it adds to be A/B measured before it can be merged.
…em (celeris#667)

chanReader decided to pause under pausedMu and applied the decision to the
engine after releasing it (engineread.go:225-232), and Read's resume branch did
the same in reverse (engineread.go:309-316). The engines' PauseRecv/ResumeRecv
closures are Swap-based and ignore the previous value -- their early return
skips only the eventfd wakeup, never the state write
(engine/iouring/worker.go:2200-2231, engine/epoll/loop.go:1672-1703) -- so
whichever callback reached the engine LAST won outright and nothing reconciled
the two. A resume could be applied before the pause it was meant to cancel,
leaving the engine's recv PAUSED while the reader believed it was RUNNING. The
reader is edge-triggered, so it never asked again: that connection stopped
delivering inbound data for the rest of its life.

Both callbacks are now invoked with pausedMu held, so the order the engine
observes equals the order of the pausedState transitions.

The fix is engine-agnostic. chanReader is the only implementation of this
decide-then-apply pair (`git grep pausedMu` finds engineread.go and tests only),
so there is no epoll-side copy to change; both engines' closures already have
the identical Swap shape and are unmodified here.

LOCK ORDER: pausedMu -> detachQMu, documented on requestPause. It cannot cycle.
pausedMu is an unexported field of an unexported type in this package and is
taken in exactly two places, both here; the engine packages do not import
middleware/websocket, so no detachQMu holder can reach either. All 24 detachQMu
critical sections in both engines (23 in engine/{iouring,epoll}, 1 in a test)
were read: every one is a straight-line queue append plus an atomic store with
no call out of the engine, and both drainDetachQueue implementations release
the lock before their loop. The eventfd write is on an EFD_NONBLOCK descriptor
and happens after detachQMu is released, so it cannot block under pausedMu.
pausedMu never nests with spillMu either: refillFromSpill returns before the
resume branch, spillChunk returns before requestPause, and hasSpill is atomic.

Measured, in Docker on linux/arm64, arms interleaved round by round, 20 samples
each, medians with a Mann-Whitney U test and an A/A control for the noise floor
(scripts and raw logs: wf-celeris-667-logs/ab, regenerate with report.sh):

  - End to end (8 WS clients, MaxBackpressureBuffer 16): epoll is unchanged at
    both memlocks (-0.5%/+0.4%, p>=0.40) and io_uring with 4 workers is
    unchanged (+0.5%, p=0.67). io_uring capped to ONE worker -- the contended
    shape -- is +16.7% ns/op / -14.4% MB/s, p=0.13 against an A/A floor of ~6%:
    suggestive, not significant.
  - Isolated chanReader: where the watermarks are crossed every few chunks the
    cost is real and significant (+45% like-for-like, p<0.002 at both memlocks);
    at the 256-chunk production default +5.4% (p=0.001) and +25.1% (p=0.0003)
    in two batches whose A/A floors were +0.2% and +2.3%.
  - The no-edge control, where the fix cannot matter, shows no fix-specific
    difference.

This is a correct-but-slower change on the contended io_uring shape and is
deliberately NOT presented as free.

The A/B also measured the bug itself: with the watermarks crossed every few
chunks, 9 of 20 (8 MiB) and 13 of 20 (128 MiB) BASE samples lost a pause/resume
mid-run and degenerated -- their edge rate collapsed ~40x. 0 of 40 fix samples
degenerated.

Adds middleware/websocket/engineread_bench_linux_test.go, the benchmarks the
A/B is built on. They report their own edge counts so the two arms can be shown
to have run the same trigger, and the no-edge control asserts that no callback
fired.
…(celeris#667)

The two celeris#667 regression tests ran in no CI job, and neither did the
other ~14 chanReader unit tests: the root race job excludes
./middleware/websocket, and the one websocket step selected
`-run '^TestBackpressure'`, which matches neither of them.

Widen that step to `^(TestBackpressure|TestChanReader)` and add the interlock
the `adaptive` (#660) and `iouring` (#664) jobs established. `-run` is a regex,
so a rename would select nothing while `go test` still exits 0; the tally
requires exactly the two celeris#667 tests to report PASS, and a skipped test
prints `--- SKIP`, never `--- PASS`, so a skip cannot satisfy it either.

Unlike those two jobs this needs no raised privileges and has no environment
skip to forbid: the tests are in-process over newChanReader, with no sockets,
no engine and no io_uring ring. `-v` is required for the tally, and the budget
goes to 300s because the step now runs two suites with per-test output; the
measured wall time of this exact shape is ~50s.

Proved in the CI shape (golang:1.27, -race, memlock 8 MiB, the step body
extracted verbatim from ci.yml and run as `bash --noprofile --norc -eo
pipefail`): fix rc=0 with passed=2; base rc=1 with both tests FAIL; both tests
renamed out of the namespace rc=1 on the interlock; both tests skipped rc=1 on
the interlock.
… drained to empty

TestChanReaderStalePauseConverges lays out the celeris#672 interleaving in one
goroutine, so the scheduler decides nothing: the crossing append's channel send
and depth snapshot, then the handler draining to empty through the real Read,
then the crossing append's requestPause applying the now-stale decision. The
oracle is the quiesced state: nothing is buffered, so the engine must not be
left paused, the reader's view must agree with the engine's, and the next
delivered chunk must reach the next Read.

It fails at every capacity swept (1, 8, 16, 256) on this commit, which already
carries the celeris#667 fix: #672 is a stale decision applied in the right
order, not two applications reordered.
…eleris#672)

requestPause applies a decision Append took from a depth snapshot before
pausedMu was held. If the handler drained to empty in between, every resume
check it made saw pausedState == false, the pause then landed on an empty
buffer, and Read, which re-evaluates the resume only after a successful
dequeue, could never lift it: the connection stopped delivering for good.

After applying the pause, requestPause now re-checks the watermark under the
same lock and resumes at once if the depth is already at or below lowWater
with nothing spilled. Once pausedState is true under the lock, every later
dequeue is followed by a resume check that sees it, so the pause cannot go
stale again.

This is arm C of the PR #671 A/B, byte-identical in code to the measured
variant (sha256 123165ca...); only its comment changed. It wedged 0/640
connections at cap16 and 0/640 at cap256 there, against 615/640 for the
celeris#667 fix alone at cap256.
… panics

Since celeris#667 both engine callbacks run with pausedMu held, and the lock
is released by a plain Unlock after them. A panic inside a callback skips the
Unlock, so when a caller recovers it the reader keeps pausedMu locked for
good: every later resume check in Read blocks, and so does the next
high-water crossing in Append, which runs on the engine worker thread.

TestChanReaderCallbackPanicReleasesPausedMu panics once inside each of the
three callback sites (the pause in requestPause, the resume in Read, the
celeris#672 re-check's resume in requestPause), recovers, and probes the lock
with TryLock, so a regression fails instead of hanging. All three fail on
this commit.
…nnot keep it

Both critical sections that apply an engine callback under pausedMu now
release it with defer: requestPause, and Read's resume branch, which moves
into resumeIfDrained so the deferred unlock ends with the decision instead of
extending the hold over Read's copy. A callback that panics no longer leaves
pausedMu held when a caller recovers, which would have blocked every later
resume check in Read and the next high-water crossing in Append on the
engine worker thread.

No behaviour change when nothing panics: the decisions, their order and the
set of statements under the lock are the same.
The #667 comments pointed at engine/iouring/worker.go:2200-2231 and
engine/epoll/loop.go:1672-1703 for the PauseRecv/ResumeRecv closures; the
rebase onto the celeris#657 work moved both, and any later engine change
would move them again. Name the closures instead, and mark the listings of
the pre-fix code as what engineread.go was at 6136554.

The benchmark's escape-hatch comment now says what the hatch is for on a
tree that carries the celeris#672 fix.
TestMeasureWedgeRate is the harness behind the PR #671 wedge numbers. Until
now it lived only with the measurement logs, so the result that decides this
PR could not be reproduced from the tree. It removes the benchmark's escape
hatch, drives each connection to its first wedge or its op budget, detects a
wedge structurally (engine paused, buffer empty, no reader progress across a
bounded number of yields) and prints one WEDGE_CONN line per connection and a
WEDGE_SUMMARY per run, with per-connection anti-vacuity checks.

The code is the measured harness unchanged; only its header comment is new:
it now says that this tree is the negative control, how to build the base and
#667-only arms from git, and the command for one observation. It skips unless
WS667_WEDGE is set, and its name stays outside ^TestChanReader so the CI step
that selects the chanReader tests never sweeps it in.
…nterlock

The websocket step's exact-PASS tally now covers four regression tests
instead of two -- the two celeris#667 orderings, the celeris#672 stale pause,
and the pausedMu panic release -- and each of the seven subtests of the last
two, so a rename, a skip, or a capacity or callback site dropped from a sweep
fails the step instead of passing it quietly. The PASS patterns end at ` (`,
never at the test name.
@FumingPower3925

FumingPower3925 commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor Author

Update: #672 now lands in this PR, and every point above is addressed

Following up on the comment above. The branch is rebased onto main 9f4d89b and now carries the celeris#672 fix. The body is rewritten, and each point is answered here with where it landed.

Also new since the last revision:

  • One mutant survives: M5, the re-check ignoring the spill. The only path I found that reaches it is a separate lost wakeup in the spill path, where a chunk spills after the handler has drained the channel and is never promoted. A single-goroutine repro leaves main 9f4d89b and this head in the same state, so it predates this PR. I am reporting it separately, not fixing it here. Correction (2026-09-27): it had not been filed when I wrote this. It is now celeris#705 (Refs middleware/websocket: a chunk that spills after the handler drained the channel is never promoted, so the connection wedges permanently (stranded spill) #705). This PR does not fix it, and mutant M5 survives because of it.
  • The whole ./middleware/websocket suite (at 8 and 128 MiB), ./engine/epoll and ./engine/iouring (at 8 and 128 MiB), base vs branch: PASS→FAIL 0 across 6 suite pairs (base vs branch, strict tally). PASS→SKIP 1: TestListenRefusesToStartWhenNoWorkerReportsItsAddress in ./engine/iouring at 8 MiB, which skips on io_uring_setup ENOMEM. That package has no diff between the arms, and in a second run of both arms main skipped it too.
  • celeris#633's cheap reproducer (-race -covermode=atomic), 10 rounds per arm: main failed 10/10 rounds and this branch 10/10 (two-sided Fisher p=1), so the failure predates this PR, and this PR does not move it at a resolution this n has.

CI on d3472e7: all 9 jobs green, with 0 FAIL and 0 SKIP result lines in the job logs. The websocket step reads want 4, passed 4; subtests want 7, passed 7.

…he callbacks

tryEngineUpgrade calls SetPauser after Detach and after the 101 is written,
on the goroutine running the upgrade, while the engine worker may already be
appending the peer's first frames. An Append that crosses highWater reads
r.pause in requestPause, and since celeris#672 r.resume in its re-check.
Nothing orders those reads with SetPauser's writes: detachQMu, the only lock
the upgrade shares with the worker, is released before SetPauser runs.

TestChanReaderSetPauserOrderedWithAppend starts the two sides together with
no synchronisation between them. Under -race it fails on this tree at both
reads (engineread.go requestPause's nil check and the re-check); without
-race it skips, since there is no wrong value to observe.
SetPauser now writes r.pause and r.resume under pausedMu, and requestPause
reads r.pause under it (the nil check moves below the Lock; the celeris#672
re-check's read of r.resume was already under it). That is the happens-before
edge the upgrade goroutine and the engine worker lacked.

Nothing changes in steady state: with the callbacks installed, requestPause
took pausedMu anyway. Only a reader with no callbacks now takes the lock
before returning. Read's reads of r.resume stay unlocked, because they run on
the handler goroutine, which the upgrade starts after SetPauser returns.
…#666 lock order)

celeris#666, which the rebase brought under this branch, made the engines'
pause/resume callbacks wake the loop through wakefd.WakeFD.Signal, which
takes a sync.RWMutex read lock. Since celeris#667 those callbacks run with
pausedMu held, so pausedMu now nests that read lock, and sync.RWMutex queues
a new reader behind a waiting writer. The lock-order note on requestPause
predated that and said nothing about it.

The note is re-derived: under pausedMu the callbacks take detachQMu and then,
after releasing it, the WakeFD read lock. Both are leaves. WakeFD's writers,
Set and Close, run only on the loop's own thread, are never reached from the
callbacks, and hold the lock only across fcntl, close(2) and atomic stores,
so a writer never waits on pausedMu, holding the write lock or queued for it.

TestChanReaderWakeFDWritersNeverWaitOnPausedMu forces the three ways a
writer and a callback holding pausedMu can meet, with the production WakeFD
and callbacks built like the engines' own: Close while a pause holds
pausedMu, Set while a resume holds it, and a pause callback queued behind a
Close that is itself queued behind another producer's in-flight write(2).
Each must complete while pausedMu is held; a watchdog turns a deadlock into
a failure.
…sured arm, 3b6c500

The header told a reader to build the #667-only arm by deleting the
celeris#672 re-check from this tree's engineread.go. That is not the arm
PR #671's wedge A/B measured: it ran 3b6c500's engineread.go (sha256
2b64141c...), which also predates the deferred unlocks and resumeIfDrained.
The recipe is now the git show that reproduces the measured arm.
…anReader interlock

The websocket step's exact-PASS interlock now requires six regression tests
and ten subtests: the four it had, TestChanReaderSetPauserOrderedWithAppend
(which the step's -race makes observable) and
TestChanReaderWakeFDWritersNeverWaitOnPausedMu with its three forced
interleavings.
@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: ecac8998-58f4-49c6-9a57-cd5873d26910

📥 Commits

Reviewing files that changed from the base of the PR and between 371c76a and 44e9714.

📒 Files selected for processing (1)
  • .github/workflows/ci.yml

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The WebSocket reader now applies pause and resume callbacks under pausedMu and rechecks stale pause decisions. The changes add race and WakeFD lock-order tests, Linux benchmarks and a wedge-measurement harness, and CI checks for named regression-test counts.

Changes

WebSocket reader pause synchronization

Layer / File(s) Summary
Pause synchronization and regression coverage
middleware/websocket/engineread.go, middleware/websocket/engineread_test.go, .github/workflows/ci.yml
SetPauser, pause application, and drained-buffer resume handling now use pausedMu. Pause application rechecks channel depth and spill state. Tests cover callback ordering, stale pauses, panic recovery, and concurrent callback installation. CI checks named test pass counts.
WakeFD lock-order interleavings
middleware/websocket/engineread_lockorder_linux_test.go
Linux tests exercise WakeFD.Close, WakeFD.Set, and WakeFD.Signal interleavings with parked callbacks. They check callback completion, eventfd writes, and final pause state.
Linux reader and echo benchmarks
middleware/websocket/engineread_bench_linux_test.go
Benchmarks measure contended reader callback, edge, wakeup, and forced-delivery rates. A no-edge control checks that callbacks do not run. The echo benchmark uses epoll and conditionally includes io_uring.
Opt-in wedge measurement
middleware/websocket/engineread_wedge_linux_test.go
The Linux opt-in test measures per-connection reader progress and pause activity, then reports aggregate wedge counts, fractions, and rates.

Priority: ⬆️ High

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Bug fix · Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to 44e97

The CI change adds checks that the named WebSocket regression tests run and pass. No issue identified here prevents merging after normal checks.

Architecture Summary

Architecture risk: 🔵 Low · up to 44e97

The change affects 1 system.

Changed systems: middleware

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — middleware (service) was modified; 5 changed files map to changed impact.

Before / after behavior

  • observed — Modified behavior in middleware/websocket/engineread.go: SetPauser now locks pausedMu while assigning the pause and resume callbacks. Comments document the callback access ordering.
  • observed — Modified behavior in middleware/websocket/engineread.go: requestPause now keeps pausedMu held through the pause callback and defers unlocking. After applying a pause, it checks whether channel depth is at or below lowWater and no spill remains; if so, it clears the paused state and invokes resume. Comments document callback ordering, lock ordering, and the stale-pause case.
  • observed — Modified behavior in middleware/websocket/engineread.go: Adds resumeIfDrained, which locks pausedMu and, when paused and channel depth is at or below lowWater, clears the paused state and calls resume. The unlock is deferred.
  • observed — Modified behavior in middleware/websocket/engineread.go: Read now delegates its drained-buffer resume check to resumeIfDrained instead of performing the state transition and callback inline.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 70.37% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 27 functions across 5 files. (1 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main changes: applying pause/resume callbacks under the deciding lock and lifting stale pauses. It matches the pull request objectives and changeset.
Linked Issues check ✅ Passed The PR satisfies the coding requirements in [#667] and [#672]. requestPause and resumeIfDrained change pausedState and invoke engine callbacks while holding pausedMu, which preserves decision …
Out of Scope Changes check ✅ Passed The changed source, regression tests, lock-order test, benchmark, wedge measurement, and CI interlock all support the pause/resume ordering, stale-pause, safety, or performance requirements in [#667] …
Full details: Docstring Coverage

Explanation

Docstring coverage is 70.37% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 27 functions across 5 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@FumingPower3925 FumingPower3925 changed the title fix(websocket): apply the pause/resume under the lock that decides them (celeris#667) fix(websocket): apply pause/resume under the lock that decides them, and lift a stale pause (celeris#667, celeris#672) Sep 27, 2026
@FumingPower3925

FumingPower3925 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Round 3: every review point, what I did, and the proof

Head 371c76a, a fast-forward of d3472e7 by five commits (one fix, three tests, one CI change). The body is rewritten from the evidence; this maps each point to it. Evidence: lane-20260926/round3/, with MANIFEST.txt mapping every file to the script that made it.

review point action proof
MAJOR 1: "reported separately" was false The stranded spill is celeris#705 (v1.6.0). The body says Refs #705, that this PR does not fix it, and that mutant M5 survives because of it. It will be fixed on top of this PR, with its repro as the failing-first test that kills M5. The round-2 reply's sentence is struck through and corrected in place. Mutants section; #705
MAJOR 2: the shipped head never measured end to end Pre-registered end-to-end A/B, hashed before the seed was drawn. Arms: main, this PR (371c76a), the #667 fix alone (3b6c500, the earlier A/B's arm) and an A/A floor. Shapes: the CI shape (memlock 8 MiB, --cpus 4) and 128 MiB. Workloads: the backpressure and echo tests (-race, CI's env) and the earlier A/B's echo rig. 192 rounds in mirrored (ABBA-style) blocks, one container per round. Failures: both pre-registered verdicts PASS. Tests 0/48 vs main 0/48; rig 1/48 vs 2/48 (one-sided p = 0.88); the A/A is quiet. Stalls (the wedge class, post hoc from each failure's line): this PR 0/48, main 0/48, #667 alone 9/48 (p = 0.0013). Corrected in round 4: Stalls (the wedge class, a POST HOC classification from each failure's own line): this PR 0/48; main's code 1/96 (main 0/48, plus 1/48 in the A/A arm, which runs main's engineread.go again: ab3-026-m8-aa.log, client 0: … i/o timeout); #667 alone 9/48. The stall p-values are post hoc: #667 alone vs main 0.0013, vs main's code (1/96) 0.0002. The pre-registered contrast is W2 round failures, #667 alone 9/48 vs main 2/48, one-sided p = 0.025. The earlier finding replicates, and this PR removes it. Every other failure, in every arm, is an io_uring engine start at 8 MiB before any upgrade (celeris#706). A usability rule excluded 3 such start failures; counting them changes nothing, and I disclose it. Per-connection classes: only clientCloseFail is ever non-zero (this PR 20 of 9,216 connections, 17 of them in one round; A/A 3). At this n it cannot be told from the floor. Throughput: not measured. The pre-registered timing matrix never started: with no container running, the idle load was 4.3–8.8 against a gate of 1.5, and #674's B1 then took the timing lock with priority. So this condition is still open, and the PR makes no throughput claim. round3/ab/: 10-PREREGISTRATION.md + .sha256, 30-REPORT.txt, 32-SENSITIVITY.txt (post hoc), timing/33-LOADGATE.txt. Body: End-to-end A/B, Throughput
MINOR: lock order stale after #666 Re-derived at 371c76a and stated in the body and in requestPause's comment. Under pausedMu the callbacks take detachQMu and then, after releasing it, WakeFD.mu for reading; both are leaves. WakeFD's writers (Set, Close) run only on the loop thread, are never reached from the callbacks, and hold the lock only across fcntl/close(2)/atomic stores, so no writer can wait on pausedMu, holding the write lock or queued for it. TestChanReaderWakeFDWritersNeverWaitOnPausedMu (651d895, Linux, in the CI interlock) forces Close and Set while a callback holds pausedMu, and the RWMutex case: a pause callback queued behind a Close that is itself queued behind an in-flight Signal. Control: with both writers made to wait on pausedMu under the write lock, all three cases fail on their watchdog instead of hanging (05-lockorder-control/). On main's engineread.go the harness's precondition fails, because main applies the callbacks outside pausedMu.
MINOR: the #623 control was vacuous Replaced with runs where the counter is non-zero and the test passes: main 9f4d89b's own #633 rounds 14 and 5 PASSED TestBackpressurePauseDoesNotCancelInflightSend/io_uring with clientCloseFail=74 and 23 of 96 (closeTimeout=0). With the #623 assertion added, those rounds must turn FAIL. Not in this PR section; 60-cover633/rounds/cover633-{14,05}-base.log
NIT: M8/M9 abort the binary The mutant table says so. Each unlocks and relocks inside a region whose unlock is deferred, so the panic test's first panic hits fatal error: sync: unlock of unlocked mutex. Re-run on 371c76a with an ABSENT column: 5 of 26 result lines never print (the panic test, its 3 subtests, and the SetPauser test); on d3472e7 it was 4 of 25. Also added M10 and M11 (SetPauser), both killed. 04-mutants.txt
NIT: title names only #667 The title now names #667 and #672. title
NIT: wedge header recipe ≠ measured arm The header now builds the #667-only arm with git show 3b6c500:middleware/websocket/engineread.go (sha256 2b64141c…, the arm the wedge A/B ran). 69bcd1c
NIT: "0 SKIP" covers only the -v steps Stated: per-test lines exist only for go test -v runs. The Unit job's root race step, its four middleware sub-module steps and the four Driver Conformance steps run without -v, so a SKIP there cannot be counted from the log. Test Plan; 30-ci/go-test-invocations.txt
NIT: SetPauser race Fixed here (small and local; no existing issue). SetPauser writes the callbacks under pausedMu, and requestPause reads r.pause under it (0a3492d). TestChanReaderSetPauserOrderedWithAppend (9122766) fails under -race on the commit before the fix (2 races: r.pause and the re-check's r.resume) and on main's engineread.go (1 race), and passes on the fix. M10 and M11 are killed by it. This changes engineread.go, so the A/B above measures 371c76a, not d3472e7. Failing-first table

Where this leaves the PR. Every point is closed except throughput, which I could not measure on this host. The throughput measurement is pre-registered and scripted (run-ab3-timing2.sh, the same arms and seeded schedule). It needs a host whose idle load is under the gate: this laptop once dasd and the interactive apps are quiet, or a cluster node. Until it runs, the review's throughput condition still blocks the merge.

Before the push (RULE 21): the whole ./middleware/websocket suite on 371c76a, Linux, -race, at the CI shape: 230 PASS / 0 FAIL / 2 SKIP (the same two skips as main), and PASS→FAIL 0 against round 2's runs. CI on 371c76a (run 36284528042): all 9 jobs and CodeQL green; the websocket step's log reads want 6, passed 6; subtests want 10, passed 10. The interlock is re-proved in the CI shape on five arms.

@FumingPower3925

FumingPower3925 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Round 4: throughput measured, stall wording corrected

The head is unchanged at 371c76a. There is no code change, so CI (run 36284528042, all green) stands. The body is updated. Evidence is in lane-20260926/round4/, and MANIFEST.txt maps every file to its script.

review point action proof
MAJOR 2: the pre-registered timing arm (W2 median ns/op) must run before merge; its load gate cannot open on this host Run. The load gate (load1 < 1.5; this host idles at 4.3–8.8) was replaced in a hashed addendum before any timing round (11-ADDENDUM.md, hashed 06:17:06Z; first round 08:45:23Z). It became one hold of the timing lock, with no container or slot at any round's start and start load logged. Everything else is as pre-registered: the arms, both shapes, the W2 command byte for byte, the 192-round seeded order and the unchanged ab3stats.py --timing. All four cells PASS. This PR vs main: epoll +1.58% and io_uring −2.26% at 8 MiB; epoll +0.58% and io_uring +2.29% at 128 MiB. None is slower than its A/A floor at p < 0.05. The one borderline cell was replicated. io_uring at 128 MiB read p = 0.059, and its declared load sensitivity read +3.37% at p = 0.019. A pre-registered 192-round replication of that cell (hashed before its seed) reads +0.40% [−0.81, +2.29], p = 0.096: PASS. What I cannot exclude: pooling both runs gives +1.40% [−0.20, +2.54]. That view is secondary and biased upward; without the blocks that overlapped foreign load (below) it is +0.60% [−0.41, +2.28]. So a cost of up to about 2.5% on io_uring with four workers remains possible. 31-THROUGHPUT.txt, 35-CONFIRM.txt, 36-CONFIRM-POOLED.txt, 12-CONFIRM.md. Body: Throughput
MAJOR 2 (the mechanism): the worker can block on pausedMu while the handler holds it across resume() Measured, as the declared secondary. The W2 rig was re-run in each round with the mutex profile on, and each parked wait was attributed to the holder's stack. The mechanism is real. In this PR the worker parks on a handler-held pausedMu about once per 440–650 ops, for 80–180 µs on average. That is about ten times as often as on main, where most such waits last a few µs. Per-round median: 18–261 ns/op for this PR vs 0 for main, one-sided p ≤ 1.6e-4 in every cell. It is small. Amortised it costs 140–375 ns per op, 0.02–0.04% of a 0.8–0.9 ms op. Ops over 10 ms are no more frequent than on main (for example 16 vs 13 and 17 vs 10 in 14,400 ops, p = 0.36 and 0.12), and per-round p99 moves +1.1% to +3.2% (p ≥ 0.19). One of 20 uncorrected tail contrasts is under 0.05: the per-round max in epoll at 128 MiB (p = 0.042). The per-round max is noisy (A/A −16%, #667 alone +128%). Corrected in round 5: the tail effect is not resolved at this n, and a small tail cost in epoll is possible. In epoll at 128 MiB every declared tail metric leans against this PR: ops at or above 10 ms 17 vs 10 (p = 0.12); per-round slow ops and ops at or above 10 ms p = 0.095; per-round max p = 0.042 (the one uncorrected contrast of 20 under 0.05); pooled p99.9 12.2 vs 1.8 ms. That p99.9 was declared, and this reply left it out. At 14,400 ops it is the 15th slowest op, so it moves by about 10 ms on a few ops. Post hoc, at 2 ms: 29 vs 13 ops, from 15 vs 9 of 24 rounds (per-round p = 0.045). io_uring shows no lean. See the round-5 reply and the body's Throughput (round5/52-TAIL.txt). 34-SECONDARY.txt, 37-WORKERSTALL-DETAIL.txt, apparatus/zz_ab4_secondary_linux_test.go
MINOR: the Summary and the reply said "main 0/48" for stalls and quoted a post hoc p without saying so Corrected. Main's code stalled in 1/96 in round 3: main 0/48, plus the A/A arm's ab3-026-m8-aa.log, client 0: … i/o timeout. The Summary, the A/B section and the round-3 reply (edited in place, struck through) now say so. They label the stall p-values post hoc (#667 alone vs main 0.0013; vs main's code 0.0002) and give the pre-registered contrast next to them: W2 round failures 9/48 vs 2/48, p = 0.025. Round 4 adds more stall counts. In the timing matrix: main 1/48, A/A 2/48, #667 alone 11/48, this PR 0/48. In the replication: 0/64, 1/64 and 0/64. So main's code stalled in 5 of 320 rounds overall, this PR in 0 of 160, and #667 alone in 20 of 96. Summary; End-to-end A/B; round-3 reply

Conduct of the run. The timing lock was held only while celeris#674's lane did not want it. The matrix waited until #674's B1 was terminated at 08:44:40Z. It paused once at a block boundary, 09:03:53–09:15:24Z, when the orchestrator asked for the lock for #674. No round was unusable or interrupted, and no process of the user or of another lane was touched. Foreign load, disclosed: another lane's native go test -race -p 2 runs overlapped the hold twice, because its host check did not gate on the timing lock (FOREIGN-LOAD-NOTE-r4h2-20260927T0914Z.txt). The overlaps were 09:14:54–09:20:01Z (timing blocks 8–9) and 10:06:12–10:06:36Z (replication block 1). The pre-registration has no rule for that, so the primaries keep those rounds. POST HOC, dropping those whole blocks changes no verdict: the matrix is PASS in all four cells, the replication reads +0.26% (p = 0.15), and the pooled view is +0.60% (38-FOREIGN-LOAD-SENSITIVITY.txt).

Where this leaves the PR. By the pre-registered rules the throughput condition is met: no cell is slower than its floor, and the replication of the borderline cell passes. The residual risk is a cost of up to about 2.5% on io_uring with four workers, which this n cannot exclude. If that cost is real, the likely source is holding pausedMu across the callbacks, which is what #667 requires. The worker stalls measured above are one visible effect of that hold, but at 0.02–0.04% of an op they are too small to account for a 2% cost on their own. Changing how the callbacks are ordered would be a separate change, not this PR.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 5: the tail in full, every number traced to a script

The head is unchanged at 371c76a. There is no code change and no new measurement, so CI stands: run 36284528042, all green. Its Unit job log was re-read today and is byte-identical to round 3's copy. The interlock reads want 6, passed 6; subtests want 10, passed 10, and the -v steps tally 138 PASS / 0 FAIL / 0 SKIP.

The body is updated. Evidence is in lane-20260926/round5/, and MANIFEST.txt maps every file to its script. Round 5's scripts only re-read round 3's and round 4's logs.

review point action proof
MINOR 1: the tail secondary was reported selectively. The declared pooled p99.9 was dropped, and "no more frequent" / "does not show up" claims more than the data show Fixed. The body now has every declared tail metric for every cell, pooled p99.9 included, and says why p99.9 jumps. At 14,400 ops it is the 15th slowest op. In epoll at 128 MiB this PR has 17 ops at or above 10 ms and main has 10, so the two p99.9 values are 12.2 ms and 1.8 ms. A round bootstrap of p99.9 runs from about 1.5 ms to 13–17 ms in every epoll arm, main's included. The wording is now yours: not resolved at this n; a small tail cost in epoll is possible. In epoll at 128 MiB every tail metric leans against this PR, and only the per-round max is under 0.05 (p = 0.042, uncorrected). Your 2 ms count is reported and labelled post hoc. It reproduces exactly: 29 vs 13 ops (A/A 12, #667 alone 34), pooled binomial p = 0.0098. Per round, the 29 come from 15 of 24 rounds against main's 9, with at most 10 in one round, and the Mann-Whitney gives p = 0.045. The pooled binomial treats ops as independent, but a parked worker delays every connection on it, so the per-round view is the fairer one. At 1 ms the pooled binomial breaks down entirely: the A/A differs from main by 17% there. 52-TAIL.txt (tail-detail.py; it asserts its declared numbers equal round4/34-SECONDARY.txt). Body: Throughput, Tail. The round-4 reply is corrected in place (struck through).
MINOR 2: some body numbers trace to no script output, and the Reproducing inventory is stale Fixed. stall-counts.py writes 51-STALLS.txt from the raw logs. It holds every stall count per run and the pooled totals (main's code 5/320, this PR 0/160, #667 alone 20/96). It gives one-sided Fisher p = 0.1303 = C(320,5)/C(480,5). It reproduces round 3's 32-SENSITIVITY.json exactly, which it asserts. It also traces one number you did not list, "#667 alone vs main's code (1/96), p = 0.0002": the file gives 0.000218. mde-confirm.py writes 53-MDE-CONFIRM.txt. ab4-secondary.py's own MDE function gives +2.0% at 64 per arm, and +3.0% at 24, which the script asserts. With 4,000 draws per point, the pre-registered rule reaches 80% power between +1.75% and +2.0%, and at +2.29% its power is 0.93. So 12-CONFIRM.md's "about 80%" was wrong, as you found. The body says so, and the hashed file is unchanged. The inventory is corrected: ab3-sensitivity.py and loadgate-summary.py are credited, the four matrix drivers are listed with their round ranges (1–29, 30–32, 33–72, 73–192), and the round-3 timing drivers are marked as never having run a round. Round 4's list gains workerstall-detail.py, foreign-sensitivity.py and assemble4.py. 51-STALLS.txt, 53-MDE-CONFIRM.txt. Body: Stalls, again, the replication's Power bullet, Reproducing

One precision on the suggested wording. epoll runs four loops in both shapes: every round of both shapes logs epoll: workers=4, and memlock limits only io_uring. So the body says "in epoll", and names 128 MiB as the shape where it leans. At 8 MiB, the same epoll setup, the declared counts lean the same way but less: 16 vs 13 ops at or above 10 ms, p = 0.36, and p99.9 is on the same knife edge. The 2 ms count shows nothing there (19 vs 18).

Also found while tracing. The W2X rig in the same timing rounds stalled in 23 of 48 #667-only rounds and in none of the others. The body now reports this as descriptive, not pooled.

Resolve the one conflict in .github/workflows/ci.yml: main renamed the
root race step to "(excluding test/ + adaptive/ + websocket +
engine/iouring)" (celeris#662 r6), and this branch added two comment
lines above that step. Keep both: this branch's comment and main's step
name. The websocket step is still the last step of the unit job, so the
comment still holds.
@FumingPower3925
FumingPower3925 merged commit bb231d5 into main Sep 27, 2026
13 checks passed
@FumingPower3925
FumingPower3925 deleted the fix/celeris-667-6136554 branch September 27, 2026 12:31
FumingPower3925 added a commit that referenced this pull request Sep 27, 2026
celeris#705: a chunk that spills after the handler has drained the channel is
never promoted. TestChanReaderLateSpillReachesParkedRead lays out the
interleaving the issue reports (Append's select finds the channel full, the
handler drains it and parks in Read, then spillChunk and requestPause run) and
requires the parked Read to get the chunk. TestChanReaderSpillPublishedAfterDrain
covers the other order: the handler drains while spillChunk is queuing the
chunk, before spillLen is published; the next Read, after the publication,
after a close, or already waiting in the window, must still get it. Its
sequential cases also require the stale-pause re-check to keep a pause while a
chunk is spilled, which the #671 review's mutant M5 does not (celeris#716,
item 1).

celeris#706: startNativeServerWithHandle never read Start's error.
TestStartNativeServerFailsFastOnStartError starts a server whose Start fails at
once (an engine type no factory knows) and requires the helper to fail with
that error, at once, not with "server not ready within timeout" after 30 s.

All three fail on main.
FumingPower3925 added a commit that referenced this pull request Sep 27, 2026
…716 item 1)

The #671 review found that no test kills mutant M5, requestPause's
celeris#672 re-check without its !r.hasSpill() term, and that #705's repro
cannot. TestChanReaderSpillPublishedAfterDrain's sequential cases (from the
test commit) require the pause to be kept while a chunk is spilled behind an
empty channel. TestChanReaderPauseRecheckHoldsWhileSpilled adds the case where
the guard is all that stands between the engine and a resume above lowWater:
the re-check runs with the channel at lowWater and a full spill behind it
(MaxBackpressureBuffer 2, BackpressureHighPct 100, BackpressureLowPct 50,
chunks delivered before SetPauser, a Read between spillChunk and
requestPause), and M5 resumes the engine with three chunks buffered against
lowWater 1. The test passes on main's engineread.go as well as on the #705
fix; it fails only without the guard.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment