Skip to content

fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) - #674

Merged
FumingPower3925 merged 32 commits into
mainfrom
fix/celeris-662-defer
Sep 27, 2026

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #662
Fixes #675

Squash-merge only. Do not merge yet: B1 has no answer (section 10). The standing order for this PR is that it does not merge until a measurement shows no regression on the default configuration. S1 measured it twice and both runs came out INCONCLUSIVE. S1b has not produced a timing result yet (section 10), and its B arm must be this head's product code.

Merge order: this PR first, then #696. #696 (the fix for celeris#691) edits ci.yml too and has to be rebased onto this PR: its new step's comment says the unit job runs ./engine/iouring without -v, which stops being true here, and its tests use startTestEngine, which this PR changes (compatible). Until #696 lands, the engine/iouring step can go red on celeris#691 at main's own rate; the #657 and #662 steps after it now run anyway (if: ${{ !cancelled() }}), so a red there cannot hide them.

Squash message. The repository squashes with the commit messages joined, and this branch carries the abandoned design's commits (8e3b3ae, 810d8c8, 10ddfbf, 7cd9481, 27c07bd). Replace the default with the proposed message at the end of this body.

History was rewritten, on purpose. On 2026-09-26 at 17:39Z this branch was force-pushed from a941102 to 1d90b5d. a941102 had 3 commits on main 985a386 and carried the design this PR abandoned. 1d90b5d has 24 commits on main 9f4d89b1, which contains celeris#657's two fixes, #681 and #687. The rebase onto #687 was registered publicly before it was done (#687, comment 5749130618; section 6). Review round 4 then added eight commits on top, fast-forward, to 7aeb2ec (section 13).

Every file:line below is at the pushed head 7aeb2ec3340b unless another sha is named. Each figure names where it comes from: a script that regenerates it from raw logs, or the report that states it. Section 12 lists which figures are regenerated by one command and which cite a report only. The evidence root, evidence/celeris-662/, is not in this repo.

TL;DR

  • The bug. While TCP_DEFER_ACCEPT is on (main's default), the kernel keeps a client that has completed its handshake but not yet sent anything out of the accept queue. When a pause closes the listener, that client's first request is answered with a reset, and the engine records nothing: no accept, no close, no error. Every adaptive switch pauses the outgoing engine (epoll: PauseAccept silently drops connections already waiting in the accept queue — their requests get EOF (8/8, deterministic), and a promotion on a GitHub runner loses ~1/3 of 2048 connections #662). So does a standalone PauseAccept (PauseAccept still drops handshake-complete idle connections for standalone engines: TCP_DEFER_ACCEPT stays on by default outside adaptive (the residual of #662, a product call) #675).
  • The fix (approach C). Steady state is unchanged: the option stays on. At a pause, each loop or worker clears the option on its own listener and keeps that listener open, accepting and serving normally, for 1.5 s, measured on the monotonic clock from that clear. Then it runs main's drain-and-close, unchanged. A resume during those 1.5 s puts the option back on the same fd. Adaptive starts the outgoing pause under e.mu and never waits for it. A tcp_synack_retries=0 guard is added. The round-3 switchPossible gate is deleted.
  • Measured.
    • Failing-first against 9f4d89b1: RESET 16/16 on both switch directions and 8/8 standalone on both engines. All pass on this branch.
    • The ramp gate that failed on 2026-09-18 now holds, as predicted. In the r6 gate (2026-09-20), this branch had 0 errors in all 12 ramp runs, while main had 1,556 read: reset (83 to 392 per ramp-only container).
    • The keep-alive stall at a one-worker io_uring revert, re-measured with the rig that first saw it (pre-registered, 80 containers): 0 of 30 containers at 8 MiB on this head, 0 of 15 at 128 MiB, while the control (the pre-rebase bytes) stalled 1 of 10. Excluded at 8 MiB by the registered rule, bound 9.5% per container, 4.9% per revert. The control was weak: this head's 0 of 30 against its 1 of 10 is p = 0.25, not significant on its own (section 11).
    • Nineteen mutations were applied; 17 were caught by a named check; m12 and m13, the two for the rebase's adaptiveTimeout merge, survive, and section 8 says why that is expected.
    • The CI interlocks were proven both ways: the three changed steps (linger, synack with the two missing cases, adaptive) and root match in 19 of 19 local cases; the engine/iouring step's good case is real CI at this head.
    • Real CI on 7aeb2ec: run 36290508684, 9 of 9 jobs green on the first attempt and on three re-runs of the same bytes (review round 4.3; merge commit 85fb140 each time), every tally at its expected count and no --- FAIL: line in any job.
  • Open. B1 (timing) has no answer, and nothing merges until it has one. The known residuals are in section 11.

1. The loss class and its mechanism

At main 9f4d89b1, both engines set TCP_DEFER_ACCEPT on every listen socket, unconditionally:

  • epoll: engine/epoll/loop.go:3118@9f4d89b1, in createListenSocket (:3087@9f4d89b1);
  • io_uring: engine/iouring/worker.go:5495@9f4d89b1, in createListenSocket (:5468@9f4d89b1).

While the option is set, a connection whose handshake has completed but whose client has sent nothing is not in the accept queue. The kernel holds it as a request socket (TCP_NEW_SYN_RECV), and accept4 answers EAGAIN. #663 added a drain, acceptQueuedOnPause (engine/epoll/loop.go:1054@9f4d89b1, engine/iouring/worker.go:2083@9f4d89b1). It accepts until EAGAIN, and then the pause closes the listen socket. That close orphans the request socket, and the client's first request is answered with a reset. The engine records no accept, no close and no error.

Every adaptive switch pauses the outgoing sub-engine, so every switch is exposed (#662). A standalone epoll or io_uring PauseAccept is exposed the same way (#675).

The kernel facts the design rests on were read at v6.17 (evidence/celeris-662/design/DECISION.md §9):

  • The listener's current defer value is read when the client's ACK arrives (net/ipv4/tcp_minisocks.c:860-866@v6.17). Clearing the option therefore stops new deferrals at once.
  • A deferred request socket is promoted into the accept queue when its SYN-ACK timer fires, about 1 s after the SYN (TCP_TIMEOUT_INIT, include/net/tcp.h:155@v6.17). Measured on loopback, promotion took 1002 to 1086 ms (DECISION.md §4.1).
  • Once the option is cleared, syn_ack_recalc takes its non-defer branch. If tcp_synack_retries=0, it expires a deferred request at the first timer instead of promoting it (net/ipv4/inet_connection_sock.c:872-875@v6.17). Red team S1 lost 256/256 that way; with TCP_SYNCNT=1 (S2) it lost 0/256 (DECISION.md §2).

The same class seen from the client side is #686. In #687's ramp measurement, 14,554 of 14,554 resets were connections the engine had never accepted. All of them happened at a promote's listener close, and the accept queue was empty at 3,233 of 3,340 closes (evidence/celeris-657/impl/pr3-r2/measure/MECHANISM.txt). #686 is closed as a duplicate of #662.

2. What this PR does: approach C, exactly

Steady state is main's. createListenSocket(addr, deferAccept) sets the option when deferAccept is true (engine/epoll/loop.go:3189,3222-3224; engine/iouring/worker.go:5594,5623-5625). Every caller passes !cfg.DisableDeferAccept, and that field defaults to false. No per-connection code path changes, and the run loops' bodies are compiled to the same instructions as 1d90b5d's (section 10).

At a pause. Each loop or worker handles the pause on its own thread, at the top of its next iteration (engine/epoll/loop.go:457-462; engine/iouring/worker.go:1063-1068), in stepAcceptPause (engine/epoll/loop.go:1097-1112; engine/iouring/worker.go:2077-2092). It has four transitions:

  1. ACTIVE to LINGERING. deferlinger.Enter (internal/deferlinger/deferlinger_linux.go:41-71):

    • clears the option on this listener (:46-52);
    • applies the synack guard if this pause asked for it (:53-62);
    • returns the deadline, Linger after the setsockopt returned (:70). Red team A3 anchored the deadline at the pause call and lost 245 connections; A4 anchored it at the clear and lost 0 of 4,028 (DECISION.md §4.1).

    The deadline is on deferlinger's own clock, Now: nanoseconds since the package's initialisation, read with time.Since, so on the monotonic clock (internal/deferlinger/deferlinger.go:84-108). It is a distinct type, Deadline, so a deadline cannot be compared with a wall-clock reading without a conversion. A step of the wall clock can therefore neither end a linger early nor keep a paused listener open after PauseAccept has returned (review round 4; section 13).

    Linger is 1500 ms (internal/deferlinger/deferlinger.go:79). It is internal, and only tests change it (:110-162). A listener that was created without the option, or a linger of 0, closes at once (deferlinger_linux.go:43-45), because nothing on it can be deferred.

  2. LINGERING, before the deadline. Nothing changes, and connections are accepted and served through the normal path:

    • epoll: the edge-triggered listener stays registered;
    • io_uring: the multishot accept stays armed, and rearmAcceptIfPending now gates only on listenFD >= 0 (engine/iouring/worker.go:4600-4605; main's gate at worker.go:4485-4490@9f4d89b1 also required !paused).
  3. LINGERING to CLOSED, at the deadline. closeListenerAfterDrain runs. It is main's drain-and-close, moved unchanged into one function (engine/epoll/loop.go:1120-1136; engine/iouring/worker.go:2104-2145). At its end it sets listenFDClosed and wakes a waiting PauseAccept (engine/epoll/loop.go:1134-1135; engine/iouring/worker.go:2143-2144).

  4. LINGERING to ACTIVE, on a resume. deferlinger.Leave puts the option back on the same fd (deferlinger_linux.go:77-87). The fd is never closed.

The tcp_synack_retries=0 guard. BeginPauseAccept reads the sysctl once, off the loop threads (internal/deferlinger/deferlinger.go:270-282). The pausing listeners get TCP_SYNCNT=1 only when that read succeeded and returned 0 (:186-189). An unreadable value never triggers the guard, because the kernel rejects TCP_SYNCNT values below 1: a guard applied on a guess could never be undone. An unreadable sysctl, and a failed clear, guard or restore, are counted every time and logged once per engine (:250-282; deferlinger_linux.go:89-114).

How long the loops sleep while lingering.

  • epoll caps epoll_wait at the time left to the deadline, including an infinite wait (engine/epoll/loop.go:515-517, :1142-1151). This is defence only: while the listener is open, the wait is already a few ms.
  • io_uring uses adaptiveTimeout = capToDeadline(sweptTimeout(), lingerUntil) while lingering, and sweptTimeout() otherwise (engine/iouring/worker.go:1864-1873, :1892-1894). sweptTimeout is fix(engine): sweep the connections a switch leaves behind, instead of waiting for each to send again (celeris#657) #687's sweep cap over baseTimeout (:1877-1888). This is the one design-level merge in the rebase (r6/rebase/00-RESOLUTIONS.md); section 8 has its two mutants.

Engine API.

  • BeginPauseAccept() starts the pause and returns (engine/epoll/engine.go:286-289; engine/iouring/engine.go:565-568). It is exported on the concrete engines only, not added to engine.AcceptController.

  • PauseAccept() is BeginPauseAccept() followed by a wait (engine/epoll/engine.go:305-339; engine/iouring/engine.go:585-619). The wait ends when any of these holds:

    • every loop has closed its listener;
    • a resume withdrew the pause;
    • every loop has exited;
    • Linger + 1 s has passed.

    So PauseAccept now blocks about 1.5 s. It blocks on a per-engine broadcast channel (internal/deferlinger/deferlinger.go:291-345) that a loop's pause close, a loop's exit and ResumeAccept close, taken before the flags are read so no wakeup is lost, with a 100 ms re-check for the one change nothing notifies (a pause landing on a listener a resume has not re-created yet). It used to poll every millisecond through the linger (section 11).

  • shutdown sets listenFDClosed and notifies (engine/epoll/loop.go:3167-3168; engine/iouring/worker.go:5528-5529), so a pause does not wait out its bound on a loop that is exiting.

Adaptive. Both pause sites in performSwitch call beginPause (adaptive/engine.go:810, :859; helper :909-917). The call is made under e.mu, so any later switch's ResumeAccept is ordered after it. Nothing waits for the linger. The round-3 switchPossible gate and its DisableDeferAccept override are deleted (27c07bd). A switch that aborts after building the lazy standby (a driver registered during the build) now also gives the fresh engine the drain a completed switch gives the engine it leaves, toward the engine that stays active (adaptive/engine.go:789-814): what it accepts during its linger moves, instead of staying on a standby no switch may make active.

Public surface. One new field, Config.DisableDeferAccept (config.go:133-143, mapped at :275), which reaches resource.Config.DisableDeferAccept (resource/config.go:70-101). It defaults to false. It is the instant, lossless opt-out: with the option off there is nothing to linger for, and the pause drains and closes at once. There is no new interface method and no new metric. The docs are rewritten for the new behaviour: Server.PauseAccept (server.go:577-598), AcceptController.PauseAccept (engine/engine.go:33-49) and README.md:50.

3. Why C, and not A or B

The full comparison is DECISION.md §2. The three candidates were scored the same way:

design correctness performance / release gate complexity / testability sum
A: exact gate (round 3) 3 2 8 13
B: clear-and-linger, L = 1.25 s 5 8 3 16
C: per-loop clear-and-linger, 1.5 s, synack guard 7 6 4 17

Why not A. A needs no clock, but it fails two of the three correctness scopes, and it fails the release gate by construction:

B and C are the same mechanism. C wins on correctness and on API surface.

  • Correctness: C has the tcp_synack_retries=0 guard and B does not. C's 1.5 s also leaves about 0.41 s of margin over the worst promotion measured on loopback (1.086 s); B's 1.25 s leaves about 0.16 s.
  • API surface: relative to main, C adds one field, DisableDeferAccept, the opt-out that round 3 had already introduced. B would add AcceptPauseLinger (whose negative value re-enables the bug), AsyncAcceptController, ErrPauseWithdrawn and four EngineMetrics fields.

C's own weaknesses were each fixed before implementation (§5 of the record): wait bounds, the guard applied on an unreadable sysctl, the positive-control rule, and existing tests that were silently hollowed out.

4. What ships, file by file

file what
internal/deferlinger/deferlinger.go (new) the linger's shared state: DefaultLinger 1.5 s (:79), the monotonic clock and Deadline (:84-108), test hooks (:110-162), SynackRetriesZero/GuardWanted (:170-189), counters (:193-243), PauseState with its once-per-engine warnings and the wait's Changed/Notify/WaitChanged (:250-359)
internal/deferlinger/deferlinger_linux.go (new) Clear/Restore/GuardSyncnt (:14-28), Enter (:41-71), Leave (:77-87), warnOnce (:89-114)
engine/epoll/loop.go lingerUntil/deferCapable (:121-130); the per-iteration step (:457-472); stepAcceptPause (:1097); closeListenerAfterDrain (:1120); lingerTimeoutMs (:1142); shutdown sets the flag and notifies (:3167-3168); createListenSocket(addr, deferAccept) (:3189)
engine/epoll/engine.go BeginPauseAccept (:286), PauseAccept's wait (:305-339), ResumeAccept notifies (:344)
engine/iouring/worker.go the same as epoll (:359-368, :1063-1077, :2077, :2104, :5528-5529, :5594), plus the adaptiveTimeout cap (:1864-1894) and the re-arm gate (:4600-4605)
engine/iouring/engine.go BeginPauseAccept (:565), PauseAccept (:585-619), ResumeAccept notifies (:624)
adaptive/engine.go beginPause at both pause sites (:810, :859, :909-917); the drain on the abort path (:811); comments corrected
adaptive/transplant.go one comment: the old engine lingers, it is not "paused", inside switchWindowHook's window
config.go, resource/config.go, server.go, engine/engine.go, README.md DisableDeferAccept and the docs (section 2)
engine tests, both packages engine/{epoll,iouring}/pause_accept_linger_linux_test.go (the same line numbers in both), with T1 …ServesDeferredIdleConnections :439, T2 …ServesArrivalsDuringLinger :643, T3 …LingerAnchoredAtClear :682, T4 …SynackRetriesZero :454, T5 TestResumeDuringLingerRestoresDeferAccept :699, T6 TestShutdownDuringLingerIsPrompt :778, and the controls that must lose: N1 …NoLingerResetsDeferredIdleConnections :446, N2 …NoClearLosesArrivals :666, N3 …NoGuardLosesAtSynackRetriesZero :461. Also pause_accept_idle_linux_test.go (the #662 witness, its no-pause control and the config test), linger_arrivals_placement_linux_test.go (TestLingerArrivalsReachTheIncomingEngine, section 7 T3) and pause_r4_linux_test.go (round 4: TestPauseAcceptWaitIsWoken in both, TestLingerTimeoutNeverBlocksPastTheDeadline in epoll)
engine/iouring/accept_rearm_sqfull_test.go T7 TestAcceptRearmedDuringLinger (:118); TestAcceptRearmRetriedAfterSQFull rewritten for the new re-arm rule (:57)
engine/iouring/ring_budget_linux_test.go (new), listen_addr_linux_test.go, driver_test.go (startTestEngine only), the io_uring linger rigs retry a start that failed only on ring ENOMEM (engine/iouring/ring_budget_linux_test.go:101 retryRingENOMEM, :145 startRingRetried662; section 7 T2)
internal/deferlinger/deferlinger_linux_test.go, deferlinger_r4_linux_test.go (new) unit tests: guard decision, deadline anchored at the clear and on the monotonic clock, clear/restore, close at once, warnings once per engine, the wait's broadcast and no-lost-wakeup order
adaptive/switch_idle_conns_linux_test.go, adaptive/linger_switch_linux_test.go, adaptive/switch_accept_churn_test.go, adaptive/reverse_transplant_test.go, adaptive/switch_abort_drain_linux_test.go (new) both switch directions keep handshake-complete idle clients; listeners keep the option across switches; the switch does not block on the linger; a flap inside the linger; linger arrivals conserved across the transplant, with no request lost; the churn storm, at linger 0 and at the default linger; the aborted switch drains its fresh standby
test/deferab/deferab_bench_test.go (new) the churn and keep-alive benchmarks B1 runs on, with the listener socket witness
.github/workflows/ci.yml ./engine/iouring leaves the root step (:121-125) for its own step (:158-180); the celeris#657 witness step (:227), the celeris#662 linger step (:346-368) and the synack=0 step (:379-402) run if: ${{ !cancelled() }}; the adaptive tally (:551-561); the ramp tests' quarantine points at celeris#708 (:452-464)

5. Failing first

Against main 9f4d89b1 (r6/gate/logs/g7*.log, regenerated with grep -HE '^\s*--- (PASS|FAIL|SKIP): ' r6/gate/logs/g7*.log). The runs used golang:1.27, memlock 128 MiB, -race -v -count=1, one container at a time. Main's bytes have not changed since (it is still 9f4d89b1), so these stand; the branch side is re-run on this head in section 6.

A row is a real failing-first witness only if main fails the claim, not just the test's premise. Four rows are; three fail on main only at their premise, because main has no linger to observe. Those three are marked, and the mutant that gives each one a genuine catch on this branch is named (section 8).

behaviour test on main 9f4d89b1 on this branch
a switch carries a client that completed its handshake but sent nothing (#662) TestSwitchKeepsHandshakedIdleConnections FAIL: outcomes=map[RESET:16], "the epoll->io_uring switch dropped 16 of 16 clients" PASS
the same on the revert …OnRevert FAIL: outcomes=map[RESET:16] PASS
a standalone pause carries them (#675) TestPauseAcceptKeepsHandshakedIdleConnections (this PR's first commit, f0c65e9, written against main's API) FAIL on both engines: AcceptCount = 0, want 8 PASS on both
… and the rig is not the cause TestPauseAcceptIdleControl, the same rig without the pause PASS on both PASS
the switch does not wait for the linger TestPerformSwitchDoesNotBlockOnLinger FAIL at its premise only: "the outgoing epoll listeners were already closed when the switch returned". Genuine catch: m7 PASS
a flap inside the linger keeps the listeners TestFlapWithinLingerKeepsListeners FAIL at its premise only: "0 of 4 epoll listeners were open" 301 ms after the promotion. Genuine catches: m3, m4, m7 PASS
the linger's arrivals are conserved across the transplant TestLingerTransplantConserves FAIL at its premise only: "the outgoing epoll accepted nothing during its linger". Genuine catch: m17 (new in round 4; the test now also fails on a lost request) PASS

Three qualifications:

  • TestAdaptiveListenersKeepDeferAcceptAcrossSwitches passes on main. It guards against the deleted round-3 gate, which main never had, so it is not a failing-first witness.
  • TestListenSocketDoesNotDeferAccept fails on main, but it asserts the abandoned first approach (8e3b3ae). It is not counted.
  • The engine-level linger tests (T1-T7, N1-N3) call BeginPauseAccept and internal/deferlinger. Neither exists on 9f4d89b1, so these tests do not compile there, and a compile error is not failing-first evidence. What stands in for it is the mutation record (section 8), where each test has a mutant that must make it lose.

Round 4's fixes, against the reviewed head 1d90b5d (round4/FF-TALLY.txt, python3 round4/tools/ff-tally.py). Each test is in this head; to run it on 1d90b5d, a staging copy that compiles there was dropped in and removed afterwards (round4/ff/, the tree clean before and after: ff/manifest.log). The abort test's copy is byte-identical; the unit tests' copies keep the test bodies and only write 1d90b5d's deadline type where it differs.

round-4 claim test on 1d90b5d on this head
the linger deadline is on the monotonic clock TestEnterDeadlineIsOnTheMonotonicClock FAIL: "the linger deadline is 1790473372210907673, in the Unix-nanosecond domain of the wall clock" PASS
each pause warning is logged once per engine TestPauseWarningsAreLoggedOncePerEngine FAIL: "two pauses with tcp_synack_retries unreadable logged 2 warnings, want 1"; "two failed clears … logged 4 warnings in all, want 2" PASS
a lingering epoll loop never blocks past its deadline TestLingerTimeoutNeverBlocksPastTheDeadline FAIL: "lingerTimeoutMs(-1) = -1 while lingering" PASS
an aborted switch drains what its fresh standby accepts TestSwitchAbortedByADriverDrainsTheFreshStandby FAIL at m128: 16 of 16 connections the standby accepted still on it; at m8: 4 of 4 PASS: m128 13 of 32 accepted during the linger, 13 detached, 13 adopted by epoll; m8 5 of 5 moved
a standalone PauseAccept is woken, not polling TestPauseAcceptWaitIsWoken (both engines) (does not compile: no wait counters) PASS; the CPU it saves is measured in section 11

6. G5, the gate that failed on 2026-09-18, and its history

G5 compares full suites, main against branch.

2026-09-18: FAILED at the pre-rebase head 42e5098 on main 985a386 (impl/ship/STOP-DECISION.txt).

2026-09-20 10:06:45Z: the prediction was registered publicly, before the rebase (#687, comment 5749130618). Verbatim:

  1. "TestRampH1Async at m128 passes on fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 rebased onto this PR";
  2. "The F-vs-S reset protocol re-run with fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 in both arms takes the unaccepted class to zero, or near it, on both."

2026-09-20 11:08Z-13:21Z: the r6 gate, rebased head 53b52f1 against main 9f4d89b1 (r6/gate/90-GATES.md; cd r6/gate && python3 tools/g5-tally.py). The order was ABBA ABBA. Each log checks that no container ran at its start; none audits starts during the run, and native host load was not recorded.

  • (1) HELD. TestRampH1Async failed 0/6 on each arm, and TestRampH1Sync 0/6 on each. The outgoing epoll held 0 of 2,048 in 36 of 36 high phases per arm.

  • (2) was not tested as registered. The F-vs-S protocol was not re-run. What the gate measured in its place is the ramp tests' own error census, main against branch (python3 r6/ship/tools/ramp-reset-census.py):

    • main: 1,556 read: reset in 13,885,363 requests;
    • branch: 0 in 13,260,792.

    That is what (2) predicts, measured by a narrower instrument.

  • The control still catches the old defect: 42e5098 failed 2/2 (995-1,043 left); 985a386 passed 2/2.

2026-09-26: re-gate on 1d90b5d, main alongside: 28 containers, full packages, both memlock shapes, counterbalanced (r6/fix/90-FIX.md). Isolation was from other containers only. The host carried other lanes' native go test load, which a docker check cannot see (round4/30-CORRECTIONS.md §6).

  • No test went from PASS to FAIL. The 14 branch containers print no --- FAIL: line.
  • Main failed TestRampH1Async (2/2 at m128) and TestRampH1Sync (1/2).
  • In the ramp-only containers (P7), main's throughput was a fraction of the r6 gate's (Sync ok = 219,754 against about 1.2M per ramp). Main's timing there is contaminated and is not cited. The branch had 0 errors in 2,818,196 requests in the same containers, in ABBA order.

2026-09-27: round 4's suites on this head (round4/SUITE-TALLY.txt, python3 round4/tools/suite-tally.py): the touched packages, full, -race -v -count=1, --cpus 4, m128 and m8. The reference is the re-gate's main and 1d90b5d containers (main's bytes are unchanged). Other lanes' containers and native processes ran beside these. So did this lane's own stall containers, beside seven of the eight runs at c12d1b8: round 4 ran its suites as a second queue alongside the stall schedule (round4/stall/STALL-CONTEXT.txt §6). Each log records its neighbours and the host load.

package 128 MiB 8 MiB
./internal/deferlinger PASS 13 PASS 13
./engine/epoll PASS 93, SKIP 7 (the same 7 as 1d90b5d) PASS 93, SKIP 7
./engine/iouring PASS 184, SKIP 2 (the synack=0 pair) c12d1b8: PASS 180, SKIP 5 (the allow-list), plus the first wake test's race. At 7aeb2ec the local runs were starved of ring memory by other lanes' containers: one run had 1 FAIL at an engine start and 92 SKIP, another rc 0 with 115 SKIP. Real CI at 7aeb2ec: green, only the 5 allowed skips
./adaptive/... PASS 95, no SKIP; the two quarantined ramps PASS here, error census total=0, every high phase epoll=0 of 2048 PASS 89, SKIP 6 (the environment skips the adaptive job enforces)
  • No test went from PASS to FAIL because of this PR. Against 1d90b5d's containers, the top-level PASS counts rise by exactly the new tests: +7 in ./internal/deferlinger, +2 in ./engine/epoll, +1 in ./engine/iouring, +2 in ./adaptive/....
  • Two FAILs, both in ./engine/iouring at 8 MiB, neither a product defect.
    • TestPauseAcceptWaitIsWoken, first version (c12d1b8). It withdrew a pause exactly three re-check periods in and raced the wait's own timer. Rewritten in b32d39f (section 13). Since then it has passed wherever its engine could start: 3 x 3 dev containers, the suites at 128 MiB, the epoll suites, and three times in real CI.
    • TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen/drain_stops_while_held, a celeris#657 test unchanged from main, at 7aeb2ec. Its engine start failed with io_uring_setup: cannot allocate memory, 58 ENOMEM lines in that container. On main, that test's helper turns a failed start into "engine did not stop within 5s". The container ran beside another lane's container. Ring memory is charged per UID (r6/gate2/90-GATE2.md, from the kernel source), and every container here runs as root, so one container's rings shrink another's 8 MiB: under the same limit, the same ring hog fit 9 rings once and 6 once (round4/logs/race-hog-*.log). The same package at the same head on a fresh GitHub runner at 8 MiB is green, with only the five allowed skips (section 9). Run again at 8 MiB on 7aeb2ec (s2b), the package passed (rc 0), but with 115 skips and 60 ENOMEM lines: other lanes' io_uring containers had most of the 8 MiB. So the local 8 MiB runs of ./engine/iouring at this head decide nothing, and real CI stands in for them.
      • This departed from the registered rule, and round 4 did not say so (round 4.3; round4/r43/10-DEVIATION-AND-CORRECTIONS.md §1). round4/00-PLAN.md §6 says a contemporaneous base re-run decides any FAIL on FINAL that base does not show. None was made (no round-4 log is a 9f4d89b run); real CI and the ENOMEM diagnosis decided it instead. The plan's no-push stop could not act either: 7aeb2ec was pushed at 03:07:33Z, before its suites (from 03:09:15Z). The conclusion stands: the test's file is main's byte for byte, base passed the subtest 4 of 4 in the re-gate (2 at 8 MiB), and CI passed it twice at 8 MiB in each of four attempts (section 9). A late base re-run here would be starved of ring memory the same way.

7. The three threads the r6 gate left open, and what was done about each

Details: r6/fix/90-FIX.md and the round-3 comment on this PR. In short:

T1: TestDriverHTTPZeroOverhead failed on the branch 2/7 (full package, m8). It is a defect on main, not exposed by this PR.

T2: the two celeris#639 tests skipped at 8 MiB, a silent coverage hole on main that this PR closes.

  • Ring memory is uncharged 12.4-22.8 ms after a ring closes (r6/gate2/90-GATE2.md). Engine starts made back to back hit ENOMEM, read it as "io_uring unavailable" and skipped. CI ran the package without -v, so nothing showed.

  • The tests now retry a start that failed only on ring ENOMEM, for up to 10 s (engine/iouring/ring_budget_linux_test.go:101):

    • the two celeris#639 tests (engine/iouring/listen_addr_linux_test.go:49,119) and startTestEngine;
    • since round 4 also the io_uring linger rigs, through startRingRetried662.

    If io_uring works and the start still fails, the test fails. A ring ENOMEM on the 4-entry probe now counts as usable, and only a positive answer is cached (engine/iouring/ring_budget_linux_test.go:81-94).

    Not retried: a start whose constructor's 1-entry capability probe fails. New reports that as "io_uring not available on this system", not as ENOMEM. It needs almost the whole budget held, which happened here only beside other lanes' io_uring containers (12 rig skips in each of two starved local 8 MiB runs); the linger step forbids skips, so in CI it would fail, not pass.

  • CI. ./engine/iouring runs in its own -v step (ci.yml:158-179). The step fails on any skip other than the five that other steps enforce, and requires both celeris#639 tests to pass by name. Real CI: celeris#639 tests want 2, ran 2, passed 2.

  • Not settled, and with no coverage consequence: why the branch's churn test ran slower in gate2.

    • T2b's timing is void: other native load on the host.
    • T2c did not replicate the gap at its registered alpha (p = 0.0898, n = 6 per arm).

T3: two interaction mutants survived TestLingerArrivalsReachTheIncomingEngine. The test gained an ORDER check (engine/epoll/linger_arrivals_placement_linux_test.go:312,429-440; the io_uring twin at engine/iouring/linger_arrivals_placement_linux_test.go:308,427-438). It reads the adopted set first and the outgoing listeners' closed flags second, with no time constant (cd r6/fix && python3 tools/t3-tally.py):

  • unmutated: PASS in all 10 iterations of each of 4 cells (epoll and io_uring, m128 and m8);
  • m10 (sweep suppressed during the linger): FAIL in all 10 of each cell, always through ORDER;
  • m11: the placement test passes, as its header says; two sweep tests catch it in both engines.

Each cell is 2 containers x -count=5, so 2 independent observations; a mutant's catch is deterministic.

8. Mutation record

m1-m11, on 1d90b5d: all caught (r6/fix/G4-TALLY.txt, cd r6/fix && python3 tools/g4-tally.py). Each was restored by cp (RESTORE_OK 24, porcelain 0 after every restore). The code they target is unchanged here, apart from Enter's return type and PauseAccept's wait.

mutant caught by
m1 no clear T1, T2, T3, T5, T6, T7
m2 deadline at Begin TestPauseAcceptLingerAnchoredAtClear, TestEnterDeadlineAnchoredAtTheClear
m3 no restore T5 and 3 adaptive tests
m4 linger 0 7 engine tests; the two switch tests; TestFlapWithinLingerKeepsListeners
m5 no guard T4 at synack 0
m6 paused in the re-arm gate T7, TestAcceptRearmRetriedAfterSQFull
m7 blocking pause in performSwitch TestPerformSwitchDoesNotBlockOnLinger ("ForceSwitch took 1.53 s"), the flap test
m8 no flag in shutdown T6 (2.20 s against 500 ms, both engines)
m9 guard on an unreadable sysctl the two guard unit tests
m10s / m10 / m11 see section 7 T3

m12-m18, round 4 (round4/G4-TALLY.txt, python3 round4/tools/g4-tally.py; applied in a worktree at c12d1b8, whose product code is this head's, and m16 again at b52bc2a, whose wake test is this head's (only driver_test.go differs); round4/mut/manifest.log: every mutant restored by cp, sha256 equal to the HEAD blob, porcelain empty after each). Negative control, unmutated: the linger tests, the deferlinger unit tests and the two new adaptive tests, rc 0, no FAIL, no SKIP.

mutant caught by the check that fired
m12: io_uring adaptiveTimeout without the linger cap survives (full ./engine/iouring, m128 and m8, rc 0) see below
m13: io_uring adaptiveTimeout with baseTimeout in place of sweptTimeout while lingering survives (full ./engine/iouring, m128 and m8, rc 0) see below
m14: linger deadline on the wall clock (Now returns time.Now().UnixNano()) TestEnterDeadlineIsOnTheMonotonicClock "the linger deadline is 1790474520131915010, in the Unix-nanosecond domain of the wall clock". The engine linger tests pass under m14, as they must: the clock is consistent, only its domain is wrong
m15: the abort path without the drain TestSwitchAbortedByADriverDrainsTheFreshStandby, m128 and m8 "18 of the 18 connections the aborted io_uring standby accepted during its linger are still on it" (m8: 7 of 7)
m16: Notify does nothing TestPauseAcceptWaitIsWoken, both engines woken=0 at the close and at the resume (first on c12d1b8's version of the test; again on the rewritten version, where all five close cycles took 100 ms, one re-check each, and all five resume cycles returned 30-71 ms after the resume, on a re-check)
m17: a connection a lingering epoll listener accepts is shut down at once (still counted, hooks still fired) TestLingerTransplantConserves, and 5 epoll linger tests "2983 churn connections ended "EOF"". Every ledger check it had before round 4 still balanced under m17 (detached 1,583 = adopted 1,583; accepts = closes = 10,116; OnConnect and OnDisconnect equal), so the new request assertion is the catch
m18: lingerTimeoutMs clamp removed TestLingerTimeoutNeverBlocksPastTheDeadline "lingerTimeoutMs(-1) = -1 while lingering"

Why m12 and m13 survive, and why no test was added for them. Review 2 expected it, from the ORDER check's 1.25 s margin. Neither removes a correctness property; each removes timeliness bounded at 100 ms. While a listener is open, baseTimeout never exceeds its 100 ms maxWait (engine/iouring/worker.go:1897-1926). So without the linger cap (m12) an idle worker closes its listener at most 100 ms after its deadline, still accepting and serving until then. Without the sweep cap while lingering (m13), the sweep's next pass over a linger arrival is at most 100 ms away instead of on its own cadence, and every linger arrival is still placed before the first close. A test tight enough to see a delay under 100 ms would be a timing test on a CI runner, and neither mutant exposes a defect, so the plan's rule (round4/00-PLAN.md §5) says report, not add.

9. CI: the interlocks, proven both ways, and the real run

Interlocks, round 4 (round4/CI-INTERLOCK.txt, cd round4 && python3 ci/classify.py). Round 4 changed the run blocks of the celeris#662 linger step (two names added) and the adaptive step (two names added), and put if: ${{ !cancelled() }} on the #657 witness, linger and synack steps. Each step was extracted verbatim from ci.yml with PyYAML (ci/00-extract-steps*.log, whose manifest records each step's if: as parsed) and run the way GitHub runs it (bash --noprofile --norc -eo pipefail, the step's own env, a container at the runner's 8 MiB).

step cases result
celeris#662 linger (unit) good; absent, renamed, skipped, subtest skipped; m4 good rc 0 (want 28, ran 28, passed 28, SKIP lines 0); every bad input rc 1, at the tally, and m4 at go test
synack=0 good; absent and subtest skipped (the two review 2 found missing); renamed, skipped; m5 good rc 0 (want 4, ran 4, passed 4); every bad input rc 1, at the tally, m5 at go test
adaptive good; absent, renamed, skipped, subtest skipped; m4 good rc 0 (want 9, ran 9, passed 9, SKIP lines 0); every bad input rc 1
engine/iouring good; another test skipped local, at b32d39f: good rc 0 (SKIP lines 5, not enforced by another step 0; celeris#639 tests ... passed 2). The other-test-skipped case went red (rc 1) but at go test, before the tally: the startTestEngine data race (fixed in b52bc2a) plus ring ENOMEM, beside two of this lane's own io_uring stall containers (128 MiB) and one other lane's container (round4/stall/STALL-CONTEXT.txt §6). At 7aeb2ec the local runs were starved of ring memory by other lanes' containers (64 ENOMEM lines, 118 skips), and the step went red, as it must when tests skip. The clean good case is real CI at 7aeb2ec: rc 0, SKIP lines 5, not enforced by another step 0; celeris#639 tests want 2, ran 2, passed 2. The step's bad inputs (absent, renamed, skipped, subtest skipped, another test skipped, m639, m639r) were proven in r6 on this byte-identical run block
root good rc 0, and ./engine/iouring is not among its packages

if: ${{ !cancelled() }} itself is GitHub's step condition, which a container cannot evaluate. What was checked: the parsed YAML carries it on the three steps (ci/00-extract-steps*.log), actionlint accepts it, and the real run below ran every step.

The r6 interlock proof (28 of 28 cases, r6/fix/CI-INTERLOCK.txt) covers the engine/iouring step's other cases; its run block is byte-identical at this head.

Real CI on 7aeb2ec. Run 36290508684 (pull_request): 9 of 9 jobs succeeded on the first attempt. The job logs were fetched with gh api …/actions/jobs/<id>/logs --allow-escape-sequences into round4/ci/github-run-36290508684/ and tallied by bash round4/tools/ci-census.sh (round4/ci/CI-CENSUS-36290508684.txt). There is no --- FAIL: line in any job. The only --- SKIP: lines are the five on the engine/iouring step's allow-list. The runner was Ubuntu 24.04, linux/amd64, kernel 6.17.0-1022-azure. The tally lines, verbatim:

engine/iouring: SKIP lines 5, not enforced by another step 0; celeris#639 tests want 2, ran 2, passed 2
celeris#657 witness and fd-lifetime tests: want 40, ran 40, passed 40, SKIP lines 0, engines 9 at workers=1: 9
celeris#657 epoll sweep tests: want 17, ran 17, passed 17, SKIP lines 0, epoll engines 2 at loops=2: 2, SWEEPIDLE lines 2 (want 2)
celeris#662 tests: want 28, ran 28, passed 28, SKIP lines 0
celeris#662 synack=0 tests: want 4, ran 4, passed 4, SKIP lines 0
celeris#657 T4: want 1, ran 1, passed 1, SKIP lines 0, RESULT lines 1, at two workers with err=0: 1
celeris#657 placement: want 9, ran 9, passed 9, SKIP lines 0, RESULT lines 9, clean 9
celeris#662 adaptive tests: want 9, ran 9, passed 9, SKIP lines 0
celeris#657 one-worker revert tests: want 2, ran 2, passed 2, SKIP lines 0, io_uring engines 2 at workers=1: 2, summaries 2 with err=0: 2

The round-4 tests there:

  • TestPauseAcceptWaitIsWoken passed three times: in the engine/iouring step at 8 MiB, and in the linger step for both engines. At 8 MiB it logged "workers=1 default linger: pause took 1.5s, woken=1 rechecks=14".
  • TestSwitchAbortedByADriverDrainsTheFreshStandby: 23 of 32 accepted during the linger, 0 left.
  • TestLingerTransplantConserves: keep-alive err=0, churn err=0.
  • The default-linger storm: 12 switches, 36 lingers ended by a resume, 12 closes, the fd count back to baseline (33), and the engine serves.
  • T2's maxAccept ranged from 40 µs to 810 µs, against its 100 ms bound.

Re-runs of the same bytes (review round 4.3). Until then, one green run was the only real-CI sample of the new strict gates (the engine/iouring skip allow-list and the adaptive tally of 9), and the local 8 MiB runs cannot stand in for it (section 6). So run 36290508684 was re-run three times (attempts 2-4, 04:53Z-05:16Z), each on the same merge commit 85fb140. All four attempts were 9 of 9 green, and each time:

  • unit had PASS 435 / FAIL 0 / SKIP 5 (the allow-list) and adaptive had 111 / 0 / 0, with the 9 tally lines identical and no DATA RACE;
  • the engine/iouring step took 155.7-156.7 s of its 300 s, and drain_stops_while_held passed twice at 8 MiB;
  • TestLingerTransplantConserves had err=0 on both counts (839-1,584 transplanted, all adopted), and the default-linger storm passed (its dial errors, logged but not asserted: 1, 0, 1, 0).

Zero failures in four runs bounds a per-run flake rate only below about 53% (one-sided 95%). A flake would show as a red run, because every tally needs its exact count (round4/r43/CI-RERUN-CENSUS.txt, bash round4/r43/tools/ci-census-rerun.sh).

Build and lint (round4/logs/g1g2/; the same containers, pins and controls as r6/fix):

  • go build and go vet passed on all five modules for linux/amd64 and linux/arm64 (10 of 10 each), and go test -c on the five packages this PR touches, both arches. gofmt -l names only test/benchcmp_ws/bench_test.go (pre-existing, section 11).
  • golangci-lint v2.13 ("0 issues" on the root, both arches, and the four middleware modules), zizmor 1.30.0 ("No findings to report") and actionlint 1.7.12 (rc 0) were clean.
  • Each tool was given a planted defect and reported it: an unchecked unix.SetsockoptInt (errcheck, LINT_RC 1), an unpinned actions/stale@v9 (unpinned-uses, rc 14), and an unquoted $(mktemp) variable in the engine/iouring step (SC2086 on line 163 of the planted copy's ci.yml, rc 1). The actionlint control first ran on a copy without .git and exited 3 ("no project was found"), which is not a detection. With an empty git init it reported the plant (logs/g1g2/g2-actionlint-poscontrol-git.log).

10. B1: no measured regression on the default configuration — NOT ANSWERED

This PR does not merge until B1 is answered by measurement. That is the maintainer's standing order. "Instrument-limited" is not an accepted answer.

  • S1 (2026-09-18, impl/s1/) measured 42e5098 against 985a386, 40 blocks per shape. Primary and replication each gave 12 EXCLUDED, 0 REGRESSION, 4 INCONCLUSIVE (ns/op on ChurnSilent and KeepAlive). No estimate reached +2%.

  • It was inconclusive because of the instrument, not the effect (s1b/ANALYSIS.md):

    • within-round noise dominated;
    • ChurnSilent showed stalls consistent with CFS quota throttling;
    • the cells needed 155-438 blocks, and S1 capped them at 40.
  • S1b (2026-09-27) stopped before its first timed round (s1b/run/). It was pre-registered against 1d90b5d. Its static witness (E2) found one more changed function than registered (the base's adaptive: forced switches leave keep-alive connections behind — async conns miss the 1.2 s flap window in every run, and a one-worker io_uring strands sync conns until their requests fail #657 sweep) and a steady-path delta of +5 instructions where +4 was registered. Both are explained statically, but the registered rule stopped the run. Its socket witness held 8/8.

  • What round 4 changes for B1. Round 4 changes product code, so B1's B arm must now be 7aeb2ec. It does not change the steady path. A per-function comparison of the compiled test binaries of 1d90b5d and this head (round4/static/STATIC-WITNESS.txt, python3 round4/static/diff.py, code addresses and global-page offsets masked) finds (*Loop).run and (*Worker).run identical instruction for instruction:

    • linux/arm64: 1588 and 1740 instructions;
    • linux/amd64: 1509 and 1688.

    The product functions that changed:

    • the pause, resume, linger and shutdown paths;
    • performSwitch (the abort drain);
    • readers of a field at a new offset, because the PauseState embedded in each engine grew.

    So S1b's steady-path expectation carries over from 1d90b5d, and its list of changed functions has to add these.

Any REGRESSION cell blocks the merge. If S1b is still inconclusive, the result goes back to the maintainer.

11. Known residuals, and the keep-alive stall

The keep-alive stall at a one-worker io_uring revert (review round 4's major; round4/stall/STALL-TALLY.txt, python3 round4/stall/analyze.py). At 42e5098 a keep-alive connection sometimes got no answer to a request it sent just after a revert (io_uring -> epoll) with one io_uring worker: 8 of 30 rounds in the D2 mechanism probe, and main stalled once too (impl/price/99-PRICE-SUMMARY.txt, S2b/D2). It was neither attributed nor excluded, and the linger raises the traffic on exactly that path (every connection the outgoing engine accepts during its linger is transplanted: 1,459 to 1,622 per revert against 28 to 32 on main). Round 4 measured it on this head's product code against main, with D2's own rig and schedule, pre-registered before the first container (round4/00-PLAN.md §3, hashed): 5k conn/s churn plus 32 keep-alives at one request per 10 ms, four forced switches per container, one container per observation, counterbalanced, memlock 8 MiB (one io_uring worker, the CI shape) and 128 MiB (four).

Arms: A = main 9f4d89b1; B = this head's product code (built at c12d1b8; every later round-4 commit changes tests only); C = the D2 binary of 42e5098 itself, sha-pinned, as the positive control. 80 containers, 2026-09-27 01:44Z-03:05Z, all 80 valid; in A and B every switch to io_uring recorded 1 worker at 8 MiB and 4 at 128 MiB.

memlock A: main B: this head C: 42e5098 (control)
8 MiB (1 io_uring worker, the CI shape) 0 of 10 0 of 30 1 of 10: 4 keep-alives at R1 (io_uring -> epoll), answered +20 ms after the switch, the next request never answered; transplant ledger balanced (1,501 detached, 1,501 adopted). D2's signature exactly
128 MiB (4 workers) 0 of 15 0 of 15 (not run)

The counts, the bounds and p = 0.25 are the registered analysis in STALL-TALLY.txt, and the churn outcomes in the last paragraph are that file's descriptive part. Every other stall figure below, including the C row's detail, is regenerated by python3 round4/stall/context.py into round4/stall/STALL-CONTEXT.txt, which was added in review round 4.2. None of those figures was registered.

By the pre-registered rule, the stall is EXCLUDED at 8 MiB on this head. The control that let the rule decide was weak. The rule asks only that C stall at least once, and it did. B stalled in 0 of 30 containers, so the one-sided 95% upper bound on B's rate is 9.5% per container, 4.9% per revert (60 reverts). At 128 MiB, 0 of 15 gives 18.1% per container, 9.5% per revert. The only comparison the plan registered is B against today's control: 0 of 30 against 1 of 10, one-sided Fisher p = 0.25, which is not significant. Main did not stall either (0 of 25).

Cross-session comparisons (not registered; STALL-CONTEXT.txt §5, with D2 recomputed from its raw rounds). C is byte-identical to the binary that stalled 8 of 30 times in D2. Against that rate, B's 0 of 30 is lower at one-sided Fisher p = 0.0023, and against today's C and D2 pooled (9 of 40), p = 0.004. Today's 1 of 10 and D2's 8 of 30 do not differ significantly (two-sided p = 0.40), so the data do not show that the control stalled less today. The cross-session figures also assume the two sessions are comparable, and the conditions below limit that.

Conditions (STALL-CONTEXT.txt §1-3, corrected in round 4.2).

  • No container ran alone: each of the 80 had at least one other container running beside it.
  • 53 of the 80 had another lane's container beside them.
  • 40 of the 80 had one of this lane's own containers beside them. The lane ran a second queue under the same two-slot lock alongside the stall schedule: its -race suites, the G4 mutants and the CI-interlock cases. While one of those ran beside a stall container, the lane held both slots itself.
  • 27 of the 80 had only this lane's own containers beside them.

Round 4 said that another lane's container ran beside all 80. It blamed the weak control on other lanes' load and did not disclose that part of that load was this lane's own. The own-lane neighbours fell evenly across the arms (8 MiB: A 6 of 10, B 16 of 30, C 6 of 10; 128 MiB: A 6 of 15, B 6 of 15), so they do not bias B against A or C. The one control stall had only another lane's container beside it.

The host's 1-min load at the containers' starts and ends was 4.6-21 (median 7.6). D2's sampler (impl/price/swdiag/sampler.log, every 5 s, 247 samples with a container running) saw only D2's own containers, at load 1.2-4.0 (median 2.3). Round 4 gave that range as 1.2-3.4, but the sampler's maximum is 3.97.

What the result bounds is this head's rate under today's conditions: below 9.5% per container at 8 MiB. It cannot exclude a rate as low as the one the control showed today. Descriptively, the churn connections that main lost at a listener close (read: reset, the #662 class) were 27 at 8 MiB and 54 at 128 MiB. This head lost 1 in all 45 containers, plus 4 resets at dial.

Inherent to closing a listen socket, and unchanged (at 42e5098, impl/price/99-PRICE-SUMMARY.txt, E3):

  • a handshake still in flight at the final close;
  • a child that completes between the drain's last EAGAIN and the close.

On loopback with arrivals every 2 ms, this branch lost 0 of 409 at the close and 0 of 72,077 during the linger. Main lost 311/311 and 125/125.

Lossy paths (E4, netem, at 42e5098). With no packet loss, 0 of 2,048 silent clients were lost at every RTT up to 400 ms. At loss p the branch loses about 1 − (1 − p)² of them, because both the retransmitted SYN-ACK and its answer must arrive: 2.0-2.4% at 1% loss. The opt-out loses about p. Main loses 100%. Not covered: non-Linux clients, middleboxes, a BPF TCP_TIMEOUT_INIT override.

The price of the pause.

  • A standalone PauseAccept now blocks about 1.5 s (epoll p95 1503.4 to 1505.0 ms; io_uring idle p95 1526 to 1536 ms, at 42e5098, S2a), and keeps admitting connections during that time.
  • A standalone PauseAccept used to spend CPU in its own wait: it slept 1 ms at a time through the linger, and on an idle io_uring engine every wakeup woke a parked thread. D1 measured 65-73 ms per pause at 42e5098. Round 4 fixed that and measured it before and after, pre-registered (round4/00-PLAN.md §4; round4/cpu/CPU-TALLY.txt, python3 round4/cpu/analyze.py), with D1's rig: 3 containers per head per shape, 4 cycles each. Per container, the median CPU of the wait (D1's c3 - c2) was, on io_uring, 43.1 to 66.5 ms on 1d90b5d and 0.8 to 4.0 ms on this head; on epoll the medians per shape went from 16.4 and 16.6 ms to 2.0 and 5.1 ms. Two of the three registered predictions held (P1, P3). P2, the pause's wall time within ±10 ms, missed at 8 MiB: +13.1 ms (median, this head minus 1d90b5d); at 128 MiB it was -4.5 ms. That wall time is set by when an idle io_uring worker notices the pause (up to 100 ms) plus the 1.5 s linger, not by the wait. On 1d90b5d alone the per-container medians at 8 MiB range from 1504 to 1523 ms.
  • After a switch, the outgoing sub-engine keeps its SO_REUSEPORT share of new connections for 1.5 s (0.50 at m128). Everything it accepts during that time is transplanted: 5,948 to 6,126 connections per promotion, against 28 to 32 on main (S2b, at 42e5098).

Other open items.

12. Reproduce

Base dir: evidence/celeris-662/.

Regenerated from raw logs by one command each.

  • Rounds up to 1d90b5d: bash r6/ship/tools/regen-all.sh writes r6/ship/REGEN.txt. It covers the G5 tallies of the r6 gate and the re-gate, G4 (m1-m11), T1, T3, CI-INTERLOCK (r6), T2c, the S1 per-cell table, the ramp census and the CI census.
    • G4, T1, T3, CI-INTERLOCK, T2c, the S1 per-cell table, the ramp census and the CI census are byte-identical to the copies saved when they ran.
    • The two G5 tallies differ only in the print order of a Python set, and in the r6 gate's T1 section, which gained phase K's containers after the tally was first saved.
    • T2b's timing is VOID (section 7, T2). t2b/tools/tally-void.py regenerates its tally with the void containers marked; t2b/T2B-TALLY.pre-void.txt keeps the tally as first published.
  • Round 4 (r6/round4/MANIFEST.txt): python3 stall/analyze.py (the stall, the registered analysis), python3 stall/context.py (the stall section's other figures, none registered: neighbours, host load, the control stall's ledger, the cross-session comparisons with D2; added in round 4.2), python3 cpu/analyze.py (PauseAccept CPU), python3 tools/suite-tally.py (the suites), python3 tools/g4-tally.py (m12-m18), python3 tools/ff-tally.py (failing-first against 1d90b5d), python3 ci/classify.py (the interlocks), python3 static/diff.py (the compiled run loops).
  • Round 4.3 (r6/round4/r43/, run from r6/round4): python3 r43/tools/ramp-baseline.py (CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708's baseline, per container, with the r6 gate and the re-gate kept apart) and bash r43/tools/ci-census-rerun.sh (every attempt of run 36290508684).

Cited from a report, not regenerated by one command here. Each report names the scripts that produced it, beside it:

  • section 1's kernel facts, promotion times (1002-1086 ms) and red-team results (S1 256/256, S2 0/256): design/DECISION.md §§2, 4.1, 9;
  • section 2's red-team anchors (A3 245 lost, A4 0 of 4,028) and section 3's design scores and r4 ab3 figures: design/DECISION.md §§2, 4.1;
  • section 7 T2's ring-uncharge latency (12.4-22.8 ms, n = 40): r6/gate2/90-GATE2.md, tallied by r6/gate2/tools/t2-tally.py;
  • section 11's residual, netem and pause-price figures at 42e5098: impl/price/99-PRICE-SUMMARY.txt, regenerated by the scripts in impl/price/tools/.

Every container log records its head sha, git status --porcelain and docker ps; round 4's also record the shared slot, the containers that started during it, and the host load.

13. Review round 4: what changed on top of 1d90b5d

Three reviews of 1d90b5d (one APPROVE, two REQUEST_CHANGES) raised two majors and a set of minors and nits. The round-4 comment on this PR answers each one. The commits, fast-forward on 1d90b5d:

30ba519 fix(engine): take the linger deadline on the monotonic clock, wake PauseAccept instead of polling, and log each pause warning once per engine
85545b2 fix(adaptive): drain what a fresh standby accepts while an aborted switch lingers it
551d62f test(engine): stop the linger rigs' goroutines on an early failure, retry their io_uring start on ring ENOMEM, and log the accept margin
71cc7a0 test(adaptive): fail the linger's conservation test on a lost request, and storm the switch with the default linger too
c12d1b8 ci: point the ramp quarantine at celeris#708, run the #657 and #662 steps after a red step, and tally the new tests
b32d39f test(engine): assert PauseAccept's wakeup over several pauses, not on one that can race its re-check
b52bc2a test(iouring): give each retried engine start its own done channel
7aeb2ec test(iouring): log when startTestEngine retried a start on ring ENOMEM

Review round 4.3 (three minor findings, no code change): #708's baseline was split into the r6 gate and the loaded re-gate (section 11); the section-6 deviations were disclosed (round4/r43/10-DEVIATION-AND-CORRECTIONS.md); and CI was re-run three times on the same bytes, all green (section 9).

Proposed squash commit message

Title:

fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) (#674)

Message:

While TCP_DEFER_ACCEPT is on, the kernel keeps a connection whose handshake
completed but which has sent nothing out of the accept queue. An accept
pause that closed the listener then orphaned it, and the client's first
request was answered with a reset that the engine never saw. Every
adaptive switch pauses the outgoing sub-engine (#662), and so does a
standalone PauseAccept (#675).

At a pause, each epoll loop and io_uring worker now clears the option on
its own listener and keeps the listener open, accepting and serving, for
1.5 s from that clear (internal/deferlinger), then runs the existing
drain-and-close. The kernel's SYN-ACK retransmit promotes every connection
deferred before the clear within that window. A resume during the linger
restores the option on the same descriptor. The deadline is on the
monotonic clock. When net.ipv4.tcp_synack_retries reads 0, the pausing
listener also gets TCP_SYNCNT=1, without which the kernel would drop the
deferred connections instead of promoting them.

Steady state is unchanged in behaviour: the option stays on, and the run
loops gain only their checks for a pending pause or linger.
BeginPauseAccept starts a pause without waiting; the adaptive engine
uses it, so a switch does not hold its lock for the linger, and an
aborted switch drains what its fresh standby accepts. PauseAccept blocks
for the linger, woken by the listeners' close or a resume.
Config.DisableDeferAccept is the instant, lossless opt-out.

CI runs ./engine/iouring in its own step with every skip accounted for,
tallies the linger, synack=0 and adaptive tests by name, and runs those
steps even after an earlier step fails. The ramp tests stay quarantined
under #708 until six GitHub-hosted runs pass.

Fixes #662
Fixes #675

Summary by CodeRabbit

  • New Features
    • Added a DisableDeferAccept option to close listening sockets immediately when pausing acceptance.
    • With the default setting, pausing acceptance on epoll and io_uring keeps listeners open for about 1.5 seconds, allowing already-connected clients and new arrivals during that window to be accepted before sockets close.
    • Adaptive engine switches now preserve connections accepted during the transition, including when a switch is rejected.
  • Documentation
    • Updated the descriptions of pause behavior, connection handling, and the immediate-close option.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Not merging: two of three reviews returned request_changes, and I re-verified the majors at 985a386

The measurement work here is good — pre-registered, counterbalanced, amended only on base-only
data, and it changed the design decision rather than rubber-stamping it. That part I would keep.
The blockers are elsewhere. I checked each of these against main myself rather than taking a
review's word.

1. The Option 2 rationale is inverted by the default engine (blocker)

// resource/config.go:14-19 @ 985a386
func defaultEngine() engine.EngineType {
	if runtime.GOOS == "linux" {
		return engine.Adaptive
	}
	return engine.Std
}

Adaptive is the default on Linux. So setting cfg.DisableDeferAccept = true on both
sub-engines (adaptive/engine.go:307) removes TCP_DEFER_ACCEPT from the path almost every user
is on, and preserves it only for users who explicitly pin Engine: Epoll or IOUring. The body
argues the opposite — that Option 2 spares users the cost — while the PR's own numbers
(+2.46–3.51 µs CPU per churned connection) now land on the default configuration, against the
standing no-major-performance-regression release gate.

The fix is probably small: the flag is set unconditionally, before the controller decides
whether an up-switch is even possible (connSwitchEnabled is computed later, and is false for
H2C or when io_uring is not viable). An adaptive server that can never switch never pauses accept
and should keep the option. Gate the flag on that decision and the default path stops paying for
a pause it will never perform.

2. The shipped tests are not the tests that were proven to fail (major)

The failing-first run was recorded against the test files committed in 985e2d2. The fix commit
82be6d1 then rewrote both of them (engine/epoll/pause_accept_idle_linux_test.go +117,
engine/iouring/… +116), and the shipped versions reference DisableDeferAccept and call
createListenSocket("127.0.0.1:0", tc.deferAccept). On main the signature is
func createListenSocket(addr string) (int, error) — one argument, at loop.go:3009 and
worker.go:5287 — so the shipped tests cannot compile against main, and the recorded "fails
on main" evidence belongs to earlier versions. Re-establish the failing-first proof with the
tests that actually ship, or state plainly which artifact proves what.

3. The engine tests can go silently absent (major)

ci.yml:127  go test -race -count=1 -timeout=300s $pkgs      # no -v

The only job that runs these packages prints no per-test lines, so a skip is invisible and the
package still reports ok. Meanwhile the io_uring rig carries unguarded t.Skipf at :142 and
:172, and the adaptive test's t.Skipf at :158 sits after requireUpSwitch, so it bypasses
CELERIS_REQUIRE_UPSWITCH entirely. This is precisely the failure mode that shipped the #656
leak, and the remedy is already in this same file — the iouring job's PASS-count interlock from
#664, ten lines away.

4. The error census is saturated, so the claims resting on it do not stand (major)

recordErr (adaptive/ramp_transplant_test.go:66-72) guards on map size, not key novelty, and
its keys embed the client's ephemeral port — so it stops counting after ~200 connections.
"READ-resets = 0" is not concludable from it. I had repeated that figure on #657 and have
withdrawn it there.

5. Closes #662 over-closes, and one tracking claim is untrue (major)

Half the loss class stays open by default: a standalone engine calling the public
Server.PauseAccept without DisableDeferAccept is still exposed. pr-body.md:192 says this
"is documented on both config fields and noted on the issue", but #662 had exactly two comments,
both predating this PR. I have now filed #675 for that residual — link it and use Refs
rather than Closes, or close #662 only once the residual is tracked and the public docs match.

6. Public documentation is left false

server.go:577 still says "PauseAccept stops accepting new connections. Existing connections
continue to be served", engine/engine.go:33 likewise, and README.md:50 similarly. After this PR
that remains untrue for standalone engines by default.

Credit where it is due

The quarantine lift was attempted, refused by CI twice on identical bytes, and reverted rather
than argued around
— with the residual correctly attributed to #657. That is the right call and
the right disclosure. The A/B's promotion of cpu_us/op to co-primary after base-only
calibration, which made significance strictly harder, is also exactly the right instinct.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 3 status: not pushed, because its own regenerated evidence does not clear the bar it set

A host reboot (an OS upgrade) wiped round 3's measurement logs mid-run. The code survived as two
local commits — 3a195c0 gates DisableDeferAccept on switchPossible(), fd77c47 documents
the ForceSwitch caveat — and every number was regenerated before anything was pushed. The
regeneration changed the conclusion, so nothing has been pushed yet.

Where no switch is reachable (memlock 8 MiB): the gated engine keeps TCP_DEFER_ACCEPT — the
socket-level witness reads exactly what main sets — and no cost is visible: cpu_us/op
−0.55% [−1.70, +0.93], +0.01% [−1.64, +1.62], −0.89% [−2.40, +0.66] on Immediate / Delayed /
Silent churn (n=30 rounds per arm, one container per round). But the same-binary floors in that
window move by up to +3.24%, with intervals reaching +17.8%, so the instrument cannot exclude
the pre-registered +2% bar. The lost round had narrated this as settled. It does not reproduce:
two regenerations both come out inconclusive, and the narrated intervals were narrower than any
regeneration in 11 of 12 cells.

Where a switch IS reachable (memlock 128 MiB — the typical host, running the default engine):
the gated fix still pays cpu_us/op +5.38% [+3.93, +6.74] (+1.77 µs per connection), +6.45%
[+4.79, +8.31] (+2.04 µs), +44.71% [+42.93, +46.67] (+7.02 µs), against tight contemporaneous
floors of +0.03%, +0.43% and −0.12%. Keep-alive is flat (+1.01%, p=0.41).

So B1 is narrowed, not answered: the default engine still carries a measured churn cost
wherever it can switch.

Also caught before pushing: round 3's new tally in the adaptive CI job ended its regex in )$,
which can never match go test's --- PASS: Name (0.01s) line. Pushed as it stood, that job would
have been red on every run, and its "failures" on bad input would have proved nothing.

Withdrawn from the current body: the error-census table (its recorder saturated at 200
distinct keys — see #657) and every figure whose evidence was lost in the reboot.

Next: before another code round, checking whether the trade-off can be dissolved instead of
chosen — keep TCP_DEFER_ACCEPT on in steady state and clear it on the outgoing listeners only
when a pause is decided, so the kernel surfaces the connections it is holding before the listener
closes. That depends on a claim about kernel behaviour, so it is being verified against the kernel
source and with a standalone probe before anyone builds on it. If it holds, it would fix standalone
Server.PauseAccept (#675) too.

…se drain (celeris#662)

#663 taught PauseAccept to serve what its accept queues hold, which fixed the
queued case. It cannot reach a connection whose handshake completed but which
has sent no data: createListenSocket sets TCP_DEFER_ACCEPT=1, so the kernel
holds that connection as a TCP_NEW_SYN_RECV request socket and keeps it out of
the accept queue entirely. The drain never sees it, the listen-socket close
orphans it, and the client's first request is answered with a reset.

Three failing-first tests, in both engines and on the adaptive switch:

  - TestPauseAcceptKeepsHandshakedIdleConnections dials 8 connections that
    write nothing, asserts the accept count BEFORE the pause -- 0 on main,
    which is the defect -- pauses, and only then sends the requests. RESET 8
    of 8 on main, on both engines and at both memlock shapes.
  - TestListenSocketDoesNotDeferAccept asserts the property on the socket
    itself. It calls createListenSocket directly rather than reading a
    loop-owned listenFD from the test goroutine, which -race would flag as a
    fault in the test instead of the engine.
  - TestSwitchKeepsHandshakedIdleConnections promotes epoll -> io_uring with
    16 such connections established: RESET 16 of 16. It is guarded by
    requireUpSwitch, so the adaptive CI job's CELERIS_REQUIRE_UPSWITCH=1 turns
    its skip into a failure.

TestPauseAcceptIdleControl is the same rig without the pause. It passes on
main, which attributes the paused arm's failure to the pause rather than to
the rig's unusual write-nothing clients. Each engine rig also logs the
kernel's own witness, TcpExtTCPDeferAcceptDrop, whose delta over the dials is
8 while the option is set and 0 once it is gone.

The 1 s deferral timer is these rigs' only clock, and it is a vacuity hazard
rather than a flake hazard: past about a second the kernel's SYN-ACK
retransmit creates the children itself and main would pass for the wrong
reason. The measured phase is bounded two orders of magnitude below that, and
the elapsed time is asserted, so an overrun fails the test loudly instead of
hiding inside it.

The setsockopt is deliberately left in place. This commit only proves the loss
class.
…eeps the clients that had not yet sent a request (celeris#662)

the other half: with TCP_DEFER_ACCEPT set, a connection whose handshake has
completed but whose client has sent nothing is never IN that queue. The
kernel keeps it as a TCP_NEW_SYN_RECV request socket and accept4 answers
EAGAIN, so acceptQueuedOnPause cannot see it, the listen-socket close
orphans it, and the client's first request is met with a reset. The engine
counts no accept, no close and no error.

Every adaptive switch pauses the outgoing engine, which is where this bites:
celeris#660's CI job saw a 2048-connection ramp lose about a third of its
connections across a promotion, every sampled error a reset.

The option is not removed outright. Pricing it first (30 rounds per arm, two
memlock shapes, counterbalanced, each cell against its own A/A floor) showed
dropping it costs +14-21% ns/op and +2.4-3.5 us of server CPU per connection
on ordinary churn, +67-200% on connect-and-never-send, and +1974% ns/op on
io_uring at the CI's one-worker memlock. Keep-alive is flat in both units
(p=0.94 and 0.20), so the cost is confined to the accept path. 19 of 24
co-primary cells regress at Holm-adjusted p<1e-8 with bootstrap CI lower
bounds far above the pre-registered +2% materiality bar, against A/A floors
that do not move.

So the option becomes configuration, defaulting to ON:

  resource.Config.DisableDeferAccept, mirrored on celeris.Config, threaded
  into both engines' createListenSocket. adaptive.New sets it on BOTH
  sub-engines, because adaptive pauses on every switch and so never gets to
  keep an option that hides connections from a pause.

An engine that never pauses keeps the throughput. A standalone epoll or
io_uring engine whose owner calls PauseAccept must set the flag; that
residual exposure is called out in the field doc and on the issue.

Tests: the two engine rigs build their engine the way one that intends to
pause must be built, and TestListenSocketDeferAcceptFollowsConfig pins both
directions of the config on the socket itself -- including that the default
still sets the option, so the measured throughput default cannot be flipped
silently. The adaptive rig needs no config of its own: if it fails, adaptive
stopped setting the flag.

CI: TestRampH1Sync and TestRampH1Async come off the `adaptive` job's -skip
list. `-skip` is a regex, so a rename would drop them silently and go test
would still exit 0; the job now tallies their PASS lines and fails if it
does not see exactly two, the same interlock the `iouring` job uses.

Prose corrected where it described the old unconditional behaviour: both
acceptQueuedOnPause doc comments now say the drain can only reach the accept
queue, epoll's "Not covered" list distinguishes what is inherent to closing
a listen socket from what is configuration, and the nodelay-inherit test
says which socket asked for the deferral.
…not their only cause

This PR's first commit lifted the celeris#662 quarantine on TestRampH1Sync and
TestRampH1Async. CI disagreed, twice, so the lift is reverted here.

Two runs on GitHub-hosted runners with the fix applied and memlock unlimited
(run 35082406457, first attempt and a re-run of identical bytes):

  attempt 1: --- FAIL: TestRampH1Sync (65.02s)   --- FAIL: TestRampH1Async (56.96s)
  attempt 2: --- FAIL: TestRampH1Sync (57.07s)   --- FAIL: TestRampH1Async (56.86s)

91 PASS / 2 FAIL / 0 SKIP both times, and those two were the only failures in
the package. Reproducible, so not the celeris#670/#673 flake situation -- which
is a different pair of tests in any case.

celeris#662's own loss class IS fixed on that runner. The issue recorded sync
phase 0 holding 1363 of 2048 conns with 685 errors, every sampled one a
"read: ... connection reset by peer". With this fix phase 0 holds 2048 of 2048
with err=0 in both variants, and the job-wide count of read-resets is 0. The 52
resets that remain are at DIAL -- a handshake still in progress when the listen
socket closes, which acceptQueuedOnPause documents as not covered and inherent
to closing a listen socket.

What still fails is later phases that only these two tests exercise:
  * async phase 4: io_uring=1316 of 2048 after 20s, err=0 across 938,818
    requests -- nothing is dropped, the transplant just does not finish (#657)
  * sync phases 3 and 5: connections stranded on the down-switch

Neither is accept-time loss, so both names go back on -skip under celeris#657,
and the PASS tally interlock is removed with them: it demanded two PASS lines
that will not appear. The job's comment now records this measurement so the
next attempt starts from it rather than repeating it.

Leaving them un-skipped was the alternative and it is worse: the Adaptive job
would be permanently red, and a job that can never be green stops being read --
which is how this coverage was lost in the first place.
…ally pause (celeris#662)

Adaptive is the DEFAULT engine on Linux (resource.defaultEngine), so setting
DisableDeferAccept on every adaptive engine put the option's measured churn
cost on the default configuration -- including the many adaptive engines that
can never switch, and so never pause, and so can never hit celeris#662.

The flag is now gated on switchPossible(): the conns-per-worker up-switch
being reachable (epoll start + io_uring viable + non-h2c), OR an io_uring
start, whose always-on error-rate revert can pause it too. An old kernel, a
memlock below one io_uring worker's rings, or Protocol H2C keeps the option
and pays nothing. connSwitchEnabled is now computed from the same predicate
so the gate and the controller cannot drift apart, and the override is logged
rather than silent.

An engine that CAN switch still pays: the pause is what loses the connection,
so only an engine that never pauses can safely keep the option.

Also in this commit:

* The engine rigs ship BOTH arms. The deferred arm asserts the mechanism the
  fix rests on -- the kernel holds a handshake-complete connection that has
  sent no data out of the accept queue -- so the two arms make opposite
  assertions about the same counter and neither can pass if the flag stops
  reaching createListenSocket. The doc comments no longer claim these files
  pass on main: they name a field main does not have and cannot compile
  there. adaptive's switch test names no new API, and is the artifact that
  fails on unmodified main.

* CI can no longer go green by skipping them. The unit job gets a PASS-count
  interlock over the eight celeris#662 engine tests (-v, memlock raised,
  CELERIS_REQUIRE_IOURING_WORKERS=1), copying the iouring job's interlock from
  celeris#664; the adaptive job gets one over the three gate tests. io_uring's
  two unguarded skips now route through skipOrFail656, and the adaptive
  standby skip -- which sat after requireUpSwitch and bypassed
  CELERIS_REQUIRE_UPSWITCH entirely -- honours it.

* The ramp error census no longer saturates. recordErr guarded on map SIZE
  while its keys embedded the client's ephemeral port, so it stopped counting
  after 200 connections, repeats of existing keys included. Counting is now
  unbounded and per class; only the verbatim exemplars are capped. Every
  by-kind figure that rested on the old map is withdrawn rather than restated.

* Public documentation is made true: Server.PauseAccept, AcceptController and
  the README said existing connections continue to be served, which is not so
  for a handshake-complete connection that has sent nothing. The residual for
  standalone engines is celeris#675.

* io_uring's acceptQueuedOnPause doc lists both inherent residual states epoll
  lists, not one; resource.Config documents the lost connect-and-never-send
  shield as a resource question, not only a throughput one.

* The A/B benchmark that priced the option is committed as test/deferab, with
  adaptive arms added, so the measurement that chose the design can be re-run
  from the repository.

Refs #662
Refs #675
…the option

switchPossible() already documented erring toward correctness on the
CELERIS_ADAPTIVE_START=iouring escape hatch. It errs the other way in exactly
one place, and leaving that to be discovered is the habit this review was
about: ForceSwitch is exported and calls performSwitch directly, bypassing the
controller the predicate reasons about, so calling it on an engine the gate
judged non-switchable does pause a sub-engine and can drop a
handshake-complete connection that has sent nothing.

No production path reaches ForceSwitch (it is documented "for testing").
Widening the gate to cover it would put the measured churn cost back on every
non-switchable adaptive engine -- the default configuration the gate exists to
protect -- to insure a method nothing in production calls.
The gate turned TCP_DEFER_ACCEPT off on both sub-engines whenever a switch
was reachable, which put the option's churn cost on the default engine of
every switch-capable host. The fix now lives in the engines' pause instead
(clear the option on the pausing listener and keep it open for a linger),
so adaptive no longer needs to decide anything about the option.

adaptive/engine.go is back to main's bytes: switchPossible, upSwitchPossible
and the DisableDeferAccept override are gone, and connSwitchEnabled is
computed exactly as main computes it. The caller's DisableDeferAccept now
reaches both sub-engines untouched. The gate's test file and its two names
in the adaptive job's tally go with it; the tally itself is rewritten with
the CI interlocks later in this series.
…pt pause (celeris#662)

A listener with TCP_DEFER_ACCEPT holds a handshake-complete connection that
has sent nothing outside the accept queue until the kernel's SYN-ACK timer
promotes it, about one second after its SYN. A pause that closes the
listener before then resets it. The engines will instead clear the option on
the pausing listener, keep it open for a linger measured from that clear,
and only then run their drain-and-close.

This package holds what the epoll and io_uring engines share for that:

- Linger (1500 ms) and three test hooks, ClearOnPause, GuardEnabled and
  ObserveDelay, all atomics read only in the pause step;
- SynackRetriesZero and GuardWanted: TCP_SYNCNT=1 is applied to a pausing
  listener only when net.ipv4.tcp_synack_retries was read successfully and
  is 0, because the kernel rejects TCP_SYNCNT below 1 and a guard applied on
  an unreadable value could never be undone;
- Clear, Restore and GuardSyncnt, and Enter/Leave, the ACTIVE->LINGERING and
  LINGERING->ACTIVE transitions, with the deadline taken after the
  setsockopt returns;
- PauseState, which BeginPauseAccept fills off the loop threads (the sysctl
  is read once per pause, there);
- process-wide counters for lingers, closes, aborts, setsockopt failures,
  guards and unreadable sysctls.

It lives under the module's internal/ rather than engine/internal/ because
the adaptive package's tests need its hooks and counters, and Go only lets
code under engine/ import engine/internal/.
…s the listener (celeris#662, celeris#675)

A pause used to close each loop's listener as soon as the loop saw it. With
TCP_DEFER_ACCEPT on, which stays the default, a client whose handshake had
completed but which had not sent its request yet was still a request socket
outside the accept queue, so the close reset it, and the engine counted
nothing.

Now each loop, on its own thread, clears the option on its own listener,
takes its deadline deferlinger.Linger (1.5 s) after that clear, and keeps
accepting and serving through the normal path until the deadline. The
kernel promotes a deferred connection about one second after its SYN, which
raises an edge on the still-registered listener like any other accept, and
anything that arrives after the clear enters the accept queue at once. At
the deadline the loop runs the unchanged drain-and-close, now in
closeListenerAfterDrain. A resume during the linger restores the option on
the same descriptor. A listener built with DisableDeferAccept, or a zero
linger, closes at once as before.

- BeginPauseAccept starts a pause and returns; PauseAccept is Begin plus a
  wait that ends when every listener is closed, when a resume withdraws the
  pause, when every loop has exited, or after Linger + 1 s.
- listenFDClosed is true exactly while paused with no listener, and loop
  shutdown now sets it, so a PauseAccept racing Shutdown returns.
- The epoll_wait timeout is capped at the time left to the deadline. While
  a loop listens its wait is already a few ms, so this is defence.
…ses the listener (celeris#662, celeris#675)

The io_uring half of the epoll change. At a pause each worker, on its own
thread, clears TCP_DEFER_ACCEPT on its own listener, takes its deadline
deferlinger.Linger (1.5 s) after that clear, and keeps its multishot accept
armed and keeps serving until the deadline; then the unchanged
cancel/drain/close runs, now in closeListenerAfterDrain. A resume during the
linger restores the option on the same descriptor, whose accept was never
cancelled. A listener built with DisableDeferAccept, or a zero linger,
closes at once as before.

- BeginPauseAccept starts a pause and returns. It promises no prompt start:
  a plain HTTP/1 worker has no eventfd poll armed and sees the flag at its
  next ring wakeup, up to 100 ms later. Nothing depends on that, because
  each listener's deadline is taken at its own clear.
- PauseAccept is Begin plus a wait that ends when every listener is closed,
  when a resume withdraws the pause, when every worker has exited, or after
  Linger + 1 s.
- rearmAcceptIfPending gates on the listener only. It used to refuse while
  paused, which would leave a lingering worker whose accept re-arm hit a
  full SQ ring deaf until its close. TestAcceptRearmRetriedAfterSQFull is
  rewritten to that rule (re-arm while the listener is open, lingering
  included; never once listenFD < 0), and TestAcceptRearmedDuringLinger
  drives the retry inside a real linger entered through the pause step.
- adaptiveTimeout is capped at the time left to the linger deadline; an idle
  worker would otherwise overshoot it by up to 100 ms.
- listenFDClosed is true exactly while paused with no listener, and worker
  shutdown now sets it, so a PauseAccept racing Shutdown returns.
…for its linger (celeris#662)

A switch used to call PauseAccept on the engine it left and wait for its
listeners to close. That pause now lingers for about 1.5 s with
TCP_DEFER_ACCEPT cleared, so the outgoing engine can serve the clients that
completed their handshake on it before the switch but had not sent their
request yet. Waiting for that inside performSwitch would hold e.mu, and
with it Metrics() and the next switch, for the whole linger.

Both pause sites in performSwitch (the abort path for a freshly built
standby, and the final pause of the old active) now call beginPause, which
uses the sub-engines' BeginPauseAccept and falls back to PauseAccept for
test fakes. It is still issued under e.mu, so a later switch's ResumeAccept
of the same engine is ordered after it, and a resume during the linger
restores the option on the same listeners.

Comment fixes: "PauseAccept itself caps its wait to 2s" was wrong, and the
freezeState release rationale no longer depends on a pause that waits.
…ts instant opt-out (celeris#662, celeris#675)

The public docs said a client that had connected but not yet sent its
request was not carried across a pause, and told callers to set
DisableDeferAccept. That is no longer true: the pause clears
TCP_DEFER_ACCEPT on each listener and keeps it open for about 1.5 s first.

- Server.PauseAccept, engine.AcceptController.PauseAccept and README: the
  pause carries such clients, blocks about 1.5 s and keeps admitting
  meanwhile; what it can still reset (a handshake in flight at the close, a
  lost or unanswered SYN-ACK retransmit); DisableDeferAccept for a pause
  that closes at once.
- Config.DisableDeferAccept and resource.Config.DisableDeferAccept: the
  mechanism, the linger, and the opt-out. No cost figures in code comments.
  The claim that a never-sending connection "never becomes a socket the
  engine owns" was wrong: the kernel promotes it about one second after its
  SYN, and the engine accepts it then.
- acceptQueuedOnPause (both engines): the drain cannot see a deferred
  connection, which is why the pause lingers before it; the remaining
  residuals.
- createListenSocket (both engines), and the engines' BeginPauseAccept /
  PauseAccept docs without internal package names. The io_uring
  stepAcceptPause/closeListenerAfterDrain block moves above
  acceptQueuedOnPause's doc comment, which it had split.
…that must lose (celeris#662, celeris#675)

pause_accept_linger_linux_test.go, the same tests in epoll and io_uring, all
on the default configuration (TCP_DEFER_ACCEPT on), none with subtests:

- T1 TestPauseAcceptServesDeferredIdleConnections: eight clients connect and
  send nothing, the pause starts inside the kernel's one-second timer, and
  they write only after every listener has closed; all eight are accepted
  during the linger and served, the pause lasts at least a second, every
  listener reads the option off during the linger, and a later dial is
  refused.
- T2 TestPauseAcceptServesArrivalsDuringLinger: a silent dial every 15 ms
  through the pause; none that connected before the close's last 60 ms is
  lost, and every one after the clear (seen on the sockets) is accepted
  within 100 ms.
- T3 TestPauseAcceptLingerAnchoredAtClear: the loops act on the pause 800 ms
  late (deferlinger.ObserveDelay); still nothing lost, so the deadline is
  measured from each listener's clear.
- T4 ...SynackRetriesZero: T1 at net.ipv4.tcp_synack_retries=0, where the
  pausing listeners read TCP_SYNCNT 1. Skips elsewhere;
  CELERIS_REQUIRE_SYNACK0=1 makes the skip a failure.
- T5 TestResumeDuringLingerRestoresDeferAccept: same SO_COOKIEs across a
  resume during the linger, the option back on, a fresh silent client hidden
  again, and PauseAccept returns at once.
- T6 TestShutdownDuringLingerIsPrompt: a shutdown during the linger releases
  PauseAccept's wait at once.
- N1 linger 0 -> RESET, N2 clear disabled -> arrivals lost, N3 guard disabled
  at synack 0 -> RESET: each passes only by observing the loss.

Every rescue test checks its premises first: hidden within 200 ms and
TCPDeferAcceptDrop moved, dials well inside the one-second timer, and
tcp_migrate_req 0 (CELERIS_REQUIRE_LINGER_PREMISES=1 turns that skip into a
failure).

Existing tests: TestPauseAcceptDeferredIdleConnectionsAreHidden becomes the
pause-free premise control; TestPauseAcceptKeepsHandshakedIdleConnections
stays as the DisableDeferAccept (instant, lossless) arm; the queued rig waits
Linger + 2 s for the close. Cost figures are gone from the test headers.
…(celeris#662)

switch_idle_conns_linux_test.go names no API main lacks, so it can be
copied onto main or onto the round-3 gate:

- TestSwitchKeepsHandshakedIdleConnections is rewritten for the new truth:
  the silent clients are HIDDEN before the switch (TCP_DEFER_ACCEPT on), the
  previous warm-up linger is waited out first, and the clients stay silent
  until every listener the outgoing engine had has closed, then write; all
  sixteen are served. It fails on main, where the outgoing listeners close
  at once. TestSwitchKeepsHandshakedIdleConnectionsOnRevert is its
  io_uring->epoll twin.
- TestAdaptiveListenersKeepDeferAcceptAcrossSwitches: socket witness
  (/proc/self/fd, SO_COOKIE, TCP_DEFER_ACCEPT) before any switch, on the
  incoming engine after a promotion, on the same epoll listeners after a
  flap back inside their linger, and after full cycles. The round-3 gate
  fails it.

These no longer need the controller's up-switch (ForceSwitch performs the
switch), so they run at the 8 MiB memlock shape too, with one io_uring
worker; the controller is frozen so only the test switches.

linger_switch_linux_test.go:

- TestPerformSwitchDoesNotBlockOnLinger: ForceSwitch under 100 ms and
  Metrics() under 50 ms while the outgoing listeners are still open.
- TestFlapWithinLingerKeepsListeners: a switch back inside the linger
  resumes the same listeners (same SO_COOKIEs) with the option back on.
- TestLingerTransplantConserves: keep-alives transplanted across a
  promotion while churn keeps landing on the lingering epoll; detached ==
  adopted, no silent-loss bucket moves, accepts == closes == hooks.

Repairs: TestAdaptiveSwitchVsAcceptChurn sets the linger to 0 and waits for
the outgoing close after each switch, and asserts from deferlinger's close
counter that listen sockets did close under the accepts (a storm of
non-blocking switches otherwise never closes one). flapScenario spaces its
flaps Linger + 0.7 s apart so each flap closes and re-creates listeners as
before.
…from the comments (celeris#662)

Every benchmark now reports, read through /proc/self/fd before the timed
region, how many LISTEN sockets the engine has on its port (listeners) and
how many of them read TCP_DEFER_ACCEPT off (listeners_defer_off). With the
gate deleted, the steady state must have the option on every listener, on
every engine including adaptive at both memlock shapes; this makes that a
per-round reading next to deferdrop/op instead of an inference.

The package comment and the adaptive arm's comments described the round-3
gate, which no longer exists.
…d run the synack=0 guard (celeris#662)

unit job:
- The celeris#662 step now runs the linger tests (T1-T3, T5, T6), their
  negative controls (N1, N2), T7 and its unit twin, and the
  DisableDeferAccept arm with its premise and socket checks, -v with
  skipping forbidden. It asserts that exactly the listed names ran and
  passed at top level (`=== RUN   Name` with nothing after it, and
  `--- PASS: Name (`), and fails on any `--- SKIP` line at any indentation,
  because a skipped subtest still lets its parent report PASS. The expected
  count is derived from the name lists.
- New step: net.ipv4.tcp_synack_retries=0 (restored on exit), running T4
  and N3 in both engines with the same tally and skipping forbidden.

adaptive job: the tally regex ended in `)$` and could never match
`--- PASS: Name (0.01s)`, so the job could never go green. It now counts
the seven celeris#662 adaptive tests the same way, and a SKIP line for any
of them fails the step.

Every step sets pipefail explicitly as well as through `shell: bash`.
…red the option (celeris#662)

TestAdaptiveListenersKeepDeferAcceptAcrossSwitches flapped back to epoll
microseconds after the promotion, before any epoll loop had acted on the
pause, so the flap withdrew a pause no loop had seen: nothing was cleared,
nothing needed restoring, and the mutation record showed the test passing
with the restore removed. It now waits until every epoll listener still
open reads TCP_DEFER_ACCEPT off (the clear) before it flaps, and logs that
state, so the flap exercises the restore on the same listeners.
…ines (celeris#662, celeris#657)

The pause keeps a listener open for about 1.5 s with TCP_DEFER_ACCEPT
cleared, so for that long the engine a switch is leaving keeps its
SO_REUSEPORT share of NEW connections. Nothing pinned where those go, and
that is exactly what failed this PR's base-vs-branch suite gate on
2026-09-18: with the linger in and celeris#657 open, an async 2048-connection
ramp left 972-1170 conns on the outgoing epoll after 20 s, err=0. Not a loss
-- placement. celeris#657 is now fixed (#681, #687) and this is the test that
holds the joint property, so the interaction cannot regress silently.

TestLingerArrivalsReachTheIncomingEngine, in engine/epoll and engine/iouring,
with an async handler and the default configuration:

  * BeginPauseAccept, then wait for the socket witness to read the option off
    every listener -- the linger has begun;
  * PHASE A: six connections arrive, are served, and go silent while NO drain
    is set. Their dispatch goroutines park with nothing owed for them (epoll:
    askAtPark returns early with no drain; io_uring: runAsyncHandler's claim
    is not reached). This is the flap case, a second switch inside the first
    one's linger;
  * StartTransplant, still inside the linger;
  * PHASE B: six more arrive, are served, and go silent with the drain set.
    They join the live set at the TAIL, past the sweep cursor -- the case the
    cycle rule exists for (celeris#657 R2 MAJOR-1);
  * then every listener closes and the engine must hold none of them.

It asserts what a client and an operator can each check: no dial refused and
every arrival accepted by THIS engine (OnConnect is the witness), every
arrival answered 200, every one adopted by the target, ActiveConnections back
to 0, and the hand-off ledger balanced (io_uring also requires
TransplantHandoffInFlight 0).

The target serves what it adopts. A target that only holds the descriptor
cannot tell a connection that was MOVED from one that was DROPPED, because
either way its client never gets an answer -- the first draft of this test
read 12 timeouts and 12 adoptions and could not say which had happened.

Both controls lose, measured in one container each, mutation restored and
sha256-verified (evidence celeris-662/r6/rebase/logs/c05..c08):

  * linger 0 (celeris#662 off): 12 of 12 dials refused, 0 accepted, 0 served,
    on both engines.
  * the sweep and the park-boundary ask removed (celeris#657 off): epoll
    adopts 0 of 12 and still holds 12; io_uring adopts 6 of 12 and still
    holds the 6 of phase A. io_uring's phase B moves without the sweep
    because its own park claims it, which is why phase A is in the test: it
    is the only shape on that engine that the sweep alone carries.

CI: the name joins the celeris#662 linger step's tally, which runs both
packages with -v, forbids every skip and asserts the exact PASS count.
…useAccept instead of polling, and log each pause warning once per engine (celeris#662)

Review round 4 of PR #674, four product findings in the pause linger:

* The linger deadline was a wall-clock reading (time.Now().UnixNano()),
  while PauseAccept's own bound is monotonic. A backward step of the wall
  clock during a linger kept a paused listener open and accepting after
  PauseAccept had returned; a forward step ended the linger early and
  brought the resets back. deferlinger now has a clock of its own, Now:
  nanoseconds since the package's initialisation, read with time.Since, so
  on the monotonic clock. Every linger deadline is taken and read through
  it, as a distinct Deadline type, so a deadline cannot be compared with a
  wall-clock reading without a conversion. TestEnterDeadlineIsOnTheMonotonicClock
  fails on 1d90b5d (its deadline is in the Unix-nanosecond domain);
  TestClockIsMonotonic pins that the epoch carries a monotonic reading.

* A standalone PauseAccept slept 1 ms at a time through the whole 1.5 s
  linger: about 1,500 wakeups per pause, 65-73 ms of CPU on an idle
  io_uring engine. It now waits on a per-engine broadcast channel
  (PauseState.Changed/Notify): a loop's or worker's pause close and its
  exit set listenFDClosed and then notify, and ResumeAccept notifies. The
  channel is taken before the flags are read, so no wakeup is lost. A
  100 ms re-check stays only for the one change that is not notified, a
  pause that lands on a listener a resume has not re-created yet. The wait
  mutex is a leaf lock. TestPauseAcceptWaitIsWoken (both engines) pins
  that the wait ends on a notification, after a close and after a resume.
  The run loops' bodies are unchanged.

* Begin warned on every pause when tcp_synack_retries was unreadable, under
  the adaptive engine's lock on every switch, and a failed clear, guard or
  restore warned per listener per pause, although the comments said
  "logged once". Each is now logged once per engine and still counted
  every time. TestPauseWarningsAreLoggedOncePerEngine fails on 1d90b5d.

* lingerTimeoutMs(-1) returned -1, so a lingering epoll loop handed an
  infinite wait would have slept through its deadline. It is capped too.
  TestLingerTimeoutNeverBlocksPastTheDeadline fails on 1d90b5d.
…itch lingers it (celeris#662)

A switch that builds the lazy standby and then finds a driver registered
during the build aborts and pauses the fresh engine. That pause lingers
for about 1.5 s with the engine in the SO_REUSEPORT group, so it accepts
its share of new connections -- about half on a four-plus-four split --
and nothing moved them: they stayed on a standby that no switch may ever
make active (review round 4 of PR #674).

The abort path now gives the fresh engine the drain a completed switch
gives the engine it leaves, toward the engine that stays active. The
linger itself stays: closing at once would reset a connection whose
handshake the fresh engine completed during the build, the celeris#662
class. The call runs under e.mu exactly where the completed switch's
applyTransplant already does.

TestSwitchAbortedByADriverDrainsTheFreshStandby fails on 1d90b5d (the
io_uring standby keeps the connections it accepted during its linger)
and passes here: 17 of 32 landed on it, 17 detached, 17 adopted by epoll.
…etry their io_uring start on ring ENOMEM, and log the accept margin (celeris#662)

Review round 4 of PR #674, test findings:

* runLingerArrivalsL662 (both engines) stopped its dialer and its
  /proc/self/fd sampler only on the success path. A t.Fatal before that
  left the dialer dialing a freed port, one every 15 ms, for the rest of
  the package run, where a later test's engine could reuse it. Both are
  now stopped by a sync.OnceFunc that the normal path and t.Cleanup call.

* The io_uring linger rigs (startLingerL662, startLingerAsyncL674,
  runPauseIdle662) sent an engine start that failed on ring ENOMEM straight
  to skipOrFail656. They now go through startRingRetried662, which retries
  such a start for as long as retryRingENOMEM does, as startTestEngine and
  the two celeris#639 tests already did.

* ioUringUsableHere cached its first answer in a sync.Once, so a probe
  that met a transient ENOMEM turned every later budget failure in the
  process into a skip. A ring ENOMEM now means io_uring works and the
  budget is held, which is what skipUnlessIOUringUnusable exists to fail
  on; only a positive answer is cached.

* T2 logs its largest dial-to-accept time next to the 100 ms bound, so its
  margin on a CI runner is visible.
…, and storm the switch with the default linger too (celeris#662)

Review round 4 of PR #674:

* TestLingerTransplantConserves logged keep-alive and churn errors but
  asserted neither, so its conservation check had never been seen to
  fail. It now fails on any keep-alive error, on any churn failure that is
  not the reset closing a listen socket inherently allows (RESET, or at
  the dial), and on more of those than two per outgoing listener. Over 66
  earlier runs keep-alive errors were 0 every time and churn resets 0-4.
  Its churn and keep-alive drivers are stopped by t.Cleanup after an early
  failure.

* TestAdaptiveSwitchVsAcceptChurn sets the linger to 0, so the storm no
  longer covered the configuration that ships.
  TestAdaptiveSwitchVsAcceptChurnDefaultLinger storms it with the default
  linger, at intervals that land both inside a linger (a resume racing the
  accepts in flight) and after its close (a re-created listener), and
  requires the fd count to settle, the port to hold exactly the active
  engine's listeners with the option on, the engine to serve, and both
  paths to have run (deferlinger Aborts and Closes).
…teps after a red step, and tally the new tests (celeris#662)

* The adaptive job's quarantine of TestRampH1Sync and TestRampH1Async
  named celeris#662 and #657; #657 is closed and this PR closes #662, and
  the workflow's rule is one open issue per quarantine. celeris#708 now
  owns it: what the lift takes (>= 6 GitHub-hosted runs of identical
  bytes) and the evidence so far. The comment no longer carries this
  branch's history.

* The celeris#657 witness step, the celeris#662 linger step and the
  synack=0 step run `if: ${{ !cancelled() }}`, so a red engine/iouring step
  (celeris#691, pre-existing on main until PR #696) cannot hide whether
  they pass.

* The linger step's tally adds TestPauseAcceptWaitIsWoken (both engines);
  the adaptive tally adds TestAdaptiveSwitchVsAcceptChurnDefaultLinger and
  TestSwitchAbortedByADriverDrainsTheFreshStandby.
… one that can race its re-check (celeris#662)

TestPauseAcceptWaitIsWoken as first written withdrew its second pause
300 ms after starting it, exactly three periods of the wait's 100 ms
re-check timer. In the round-4 suite at 8 MiB (./engine/iouring, one
worker) the timer fired first, the wait returned on it, and the test
failed with "the resume did not wake the waiting PauseAccept" although the
resume had notified. A single pause's wake source is a race whenever its
close or resume lands on a re-check, so the test now asserts the
notification over several pauses, where only its total absence fails:

* one pause at the default linger: it takes the linger, closes every
  listener and re-checks no more often than its length allows;
* five pauses at a 50 ms linger: the listeners' close wakes the wait;
* five pauses withdrawn at 130-330 ms, never a multiple of the re-check:
  the resume wakes the wait.

With notifications removed every one of them ends on the timer, so the
mutant that disables Notify is still caught every time.

TestNotifyWakesEveryWaiter now checks the notified channels directly: its
waiters used WaitChanged, which returns true on the 100 ms re-check too,
so it passed without a notification. One message in
TestPauseWarningsAreLoggedOncePerEngine said "logged nothing" where the
count was only wrong.
…eleris#662)

startTestEngine's ENOMEM retry (9a01f71) started each attempt's Listen in
a goroutine that closed the shared `done` variable after sending Listen's
result. When an attempt failed and the next one began, nothing ordered
that goroutine's read of `done` before the next attempt's write, and
-race reported it. Round 4's iouringpkg-otherskip interlock case (the
engine/iouring step run as CI runs it, at 8 MiB, with another lane's
container running beside it) retried starts and got seven DATA RACE
reports, each failing the test that made the start. The attempt now
captures its own engine and channel.
…M (celeris#662)

The other retry sites (the two celeris#639 tests and the linger rigs'
startRingRetried662) say when a start was retried; startTestEngine did
not, so a run could not show whether its retry path, where the race fixed
in the previous commit lived, had been exercised at all.
@FumingPower3925

FumingPower3925 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Review round 4: every finding, what was done, and the proof

Head 1d90b5d → 7aeb2ec, fast-forward, eight commits. Evidence: evidence/celeris-662/r6/round4/ (MANIFEST.txt; the plan and both pre-registrations were hashed before the first container, 00-PLAN.sha256). Main is still 9f4d89b1. Merge order: this PR, then #696. B1 still holds the merge (body section 10).

The two majors

Review 1 (major): the keep-alive stall at a one-worker io_uring revert, re-measured on this head. I re-ran D2's own rig and schedule (5k conn/s of churn and 32 keep-alives, four forced switches per container). The runs used one container per observation, counterbalanced, pre-registered and hashed before the first container (00-PLAN.md §3). This head's product code was compared with main, and the 42e5098 binary that stalled in D2 served as the positive control.

memlock main 9f4d89b1 this head control 42e5098
8 MiB, one io_uring worker (the CI shape) 0/10 0/30 1/10 (4 keep-alives lost at R1, D2's signature exactly)
128 MiB, four workers 0/15 0/15 —

Under the pre-registered rule, the stall is EXCLUDED at 8 MiB, but the control was weak. The rule asks only that the control stall at least once, and it did. This head did not stall in 30 containers. The one-sided 95% upper bound is 9.5% per container, 4.9% per revert. At 128 MiB the bound is 18.1% per container. The only registered comparison is this head against today's control, 0/30 against 1/10, and it is not significant on its own (p = 0.25).

Not registered, across sessions (round4/stall/STALL-CONTEXT.txt §5): against 42e5098's D2 rate (8/30), this head's 0/30 gives p = 0.0023. Today's 1/10 and D2's 8/30 do not differ significantly (two-sided p = 0.40), so the data do not show that the control stalled less today.

Corrected in review round 4.2: this comment first said that another lane's container ran beside every one of the 80 containers. That was wrong. None of the 80 ran alone:

  • 53 of the 80 had another lane's container beside them;
  • 40 had one of this lane's own containers beside them (its -race suites, G4 mutants and CI-interlock cases, run as a second queue alongside the stall schedule);
  • 27 had only this lane's own containers beside them.

The own-lane load fell evenly across the arms, so it does not bias the comparison. The host was loaded (1-min load 4.6-21).

The result bounds this head below 9.5% per container. It cannot exclude a rate as low as the one the control showed today.

Review 3 (major): nothing would own the ramp quarantine after merge. I searched first: no open issue covered it; the only hit for TestRampH1 was #662. I filed #708 (v1.6.0). It records what the lift takes (at least 6 GitHub-hosted runs of identical bytes, then a PASS tally for both names), what failed before (run 35082406457: #657's placement; main 9f4d89b1: the #686 read resets, 6,615 over 8 ramp runs), and what the fix shows locally. The ci.yml quarantine comment now points at #708 and carries no branch history (ci.yml:452-464).

Every finding

# finding action proof
R1-0 MAJOR: the one-worker io_uring revert stall not re-measured measured on this head against main, with D2's own rig, pre-registered (00-PLAN.md §3) above; stall/STALL-TALLY.txt
R3-0 MAJOR: no open issue owns the ramp quarantine after merge searched, filed #708 (v1.6.0), ci.yml points at it above; ci.yml:452-464
R1-1 minor: linger deadline on the wall clock deferlinger.Now (monotonic) and a Deadline type every deadline goes through TestEnterDeadlineIsOnTheMonotonicClock FAILS on 1d90b5d, PASSES here; mutant m14 (wall clock) caught by it
R1-2 minor: switch-abort path lingers the fresh standby with no drain the abort path installs the drain toward the engine that stays active (adaptive/engine.go:811); the linger stays, since closing at once would reset its deferred handshakes TestSwitchAbortedByADriverDrainsTheFreshStandby FAILS on 1d90b5d (m128 16/16, m8 4/4 left on the standby), PASSES here (m128 13/13, m8 5/5 moved); mutant m15 caught at both shapes
R1-3, R2-7, R3-1 minor: #691/#696 interplay if: ${{ !cancelled() }} on the #657 witness, #662 linger and synack steps; merge order #674 then #696 stated in the body, with what #696 must rebase ci/00-extract-steps.log (if read from the parsed YAML); actionlint rc 0 (and its planted control reported); real CI run 36290508684: every step ran, 9/9 jobs green
R1-4 nit: log volume, wrong "logged once" once per engine per kind, still counted every time TestPauseWarningsAreLoggedOncePerEngine FAILS on 1d90b5d (2 warnings for 2 pauses; 4 for 2 clears), PASSES here
R1-5 nit: lingerTimeoutMs(-1) returns -1 capped TestLingerTimeoutNeverBlocksPastTheDeadline FAILS on 1d90b5d, PASSES here; mutant m18 caught
R1-6 nit: standalone io_uring PauseAccept CPU in a 1 ms sleep loop measured, then fixed: the wait blocks on a per-engine notification (close, exit, resume) with a 100 ms re-check the wait's CPU per pause, io_uring, per-container medians: 43.1-66.5 ms on 1d90b5d, 0.8-4.0 ms here (pre-registered P1 held; P3 held; P2, pause wall within ±10 ms, missed at 8 MiB by +13.1 ms, the io_uring observe latency's spread); TestPauseAcceptWaitIsWoken pins it; m16 caught
R1-7 nit: stale comment, 00-RESOLUTIONS credit, squash adaptive/transplant.go comment fixed; the credit corrected in round4/30-CORRECTIONS.md §1; a proposed squash message ends the body commit 85545b2; body
R2-0 minor: default squash writes the abandoned design proposed squash title and message at the end of the body body
R2-1 minor: quarantine comment with branch history, no open issue rewritten, points at #708 as R3-0
R2-2 minor: three failing-first rows fail on main only at their premise relabelled as premise-only; the genuine catches named (m7; m3/m4/m7; m17) body §5; m17 below
R2-3 minor: the rebase's adaptiveTimeout merge has no mutant m12 (no linger cap) and m13 (no sweep cap while lingering) run on both io_uring shapes both survive (full ./engine/iouring, m128 and m8). Each removes at most 100 ms of timeliness, which is baseTimeout's cap while a listener is open, and no correctness; the ORDER margin is 1.25 s. Body §8 explains; no test added, per the plan's rule
R2-4 minor: "every number is regenerated" too broad the preamble now says each figure names its script or its report; §12 lists which are regenerated by one command and which cite a report body preamble, §12
R2-5 minor: gate dated 2026-09-26; prediction (2) paraphrased dated 2026-09-20 (11:08Z-13:21Z); (2) quoted verbatim, and the body says it was not tested as registered: the ramp error census stood in for the F-vs-S protocol body §6; 30-CORRECTIONS.md §3
R2-6 minor: re-gate isolation overstated; P7 main-arm timing "isolated from other containers only; the host carried native load"; P7's main-arm results marked contaminated and no longer cited body §6; 30-CORRECTIONS.md §6
R2-8 nit: T3 "10/10" "2 containers x -count=5 per cell" body §7; 30-CORRECTIONS.md §5
R2-9 nit: g5-tally SKIP->FAIL blind spot this round's tally flags any FINAL FAIL that main does not also fail, a base SKIP included (tools/suite-tally.py); r6's tallies left as run, noted 30-CORRECTIONS.md §4
R2-10 nit: synack interlock lacks absent/subskip both run, with every other case the changed steps need synack-absent: want 4, ran 3 rc 1; synack-subskip: SKIP lines 1 rc 1
R2-11, R3-5b nit: ioUringUsableHere caches a transient ENOMEM a ring ENOMEM counts as usable (the budget is held, which is a FAIL); only a positive answer is cached ring_budget_linux_test.go:81-94; ./engine/iouring PASS locally at m128 and at 8 MiB in real CI (the local 8 MiB runs were starved by other lanes, body §6)
R2-12 nit: CodeRabbit !cancelled() done as R1-3
R2-13 nit: two ship agents; hold only as text the ship-state record corrected (30-CORRECTIONS.md §2). The hold stays in the body's first line: there is no blocking label in this repo, and converting to draft would change which workflows run on the PR —
R3-2 minor: two pre-existing silent CI holes unfiled searched (#684 covers ./adaptive only), filed #709 (v1.6.0): the epoll sendfile e2e tests (Workers: 1) and the #592 io_uring subtests (8 MiB) #709
R3-3 minor: goroutines leak on early failure runLingerArrivalsL662 (both engines) stops its dialer and sampler through sync.OnceFunc from t.Cleanup too; TestLingerTransplantConserves likewise for its churn and keep-alive drivers commits 551d62f, 71cc7a0
R3-4 minor: the switch storm no longer covers the default config TestAdaptiveSwitchVsAcceptChurnDefaultLinger: default linger, switches inside and after lingers; fd count, listeners, serves, and both paths taken passes locally at m128 (12 switches, 35-36 lingers ended by a resume, 12 closes, fd back to baseline), at m8, and in CI
R3-5a nit: linger rigs' engine start has no ENOMEM retry startLingerL662, startLingerAsyncL674, runPauseIdle662 go through startRingRetried662 the rigs PASS locally at m128 and at 8 MiB in real CI (the linger step forbids skips). The retry covers a start whose rings fail on ENOMEM; the constructor's 1-entry capability probe, which reports ENOMEM as "io_uring not available", is not retried. It fails only when the budget is almost wholly held, as other lanes' containers held it here (12 such skips in each of the two starved 8 MiB runs, s2, s2b); in CI the step would then go red, not pass
R3-6 nit: T2's accept margin not logged; flap warm-up margin T2 logs maxAccept next to its 100 ms bound (CI: 40-810 µs). The flap test's 1.1 s warm-up is left: if a worker has not parked, the test waits up to 1.2 s for io_uring's listeners and otherwise FAILS its premise loudly; it cannot pass falsely suite logs
R1 rec 2 TestLingerTransplantConserves only logs errors fails on any keep-alive error, any churn failure other than the listen-close reset, or more than two such resets per outgoing listener; m17 now catches it under m17 the old ledger checks all balanced (detached 1,583 = adopted 1,583, accepts = closes = 10,116), and the new check failed: "2983 churn connections ended EOF"

Re-gate on 7aeb2ec

  • Failing-first against 1d90b5d (FF-TALLY.txt). All four behavioural fixes FAIL there and PASS here: the monotonic deadline, warnings once per engine, the clamp, and the abort drain (m128 16/16 and m8 4/4 left on the standby).
  • Mutants m12-m18 (G4-TALLY.txt). m14-m18 were each caught by the predicted test. m12 and m13 survive, as expected. Every mutant was restored by cp, with sha256 equal to the HEAD blob.
  • Suites of the touched packages at m8 and m128 (SUITE-TALLY.txt). No PASS->FAIL comes from the PR. No new skip, except in the local 8 MiB ./engine/iouring runs at 7aeb2ec, where every new skip is an engine start that met ring ENOMEM.
    • The first version of the new wake test raced its own 100 ms re-check. It was rewritten (b32d39f).
    • A celeris#657 test unchanged from main failed its engine start on ring ENOMEM, beside another lane's container. The same package is green at 8 MiB in real CI. Later local runs of that package at 8 MiB were starved of ring memory (115 skips) and decide nothing.
  • Found and fixed by this round's gates. startTestEngine's retry (from 9a01f71) shared a done channel across attempts. That is a data race under -race whenever a start is retried: 7 of them in one CI-shape case. Fixed in b52bc2a. A ring hog that forces the retry showed it both ways: before, a DATA RACE and a FAIL; after, 199 retries, no race, PASS.
  • Interlocks (CI-INTERLOCK.txt, extracted verbatim from ci.yml). The linger step's 6 cases, the synack step's 6 (the two missing ones included), the adaptive step's 6 and root all match. The engine/iouring step's good case is real CI at this head. Locally at 8 MiB, other lanes' io_uring containers starved it, and it went red as designed. Its bad cases are r6's, on the byte-identical run block.
  • Real CI, run 36290508684 on 7aeb2ec: 9/9 jobs green on the first attempt, with every tally at its count (celeris#662 tests: want 28 ... passed 28, synack=0 ... 4, adaptive ... 9, engine/iouring: SKIP lines 5, not enforced by another step 0).
  • G1/G2. Build and vet passed on both arches. golangci-lint, zizmor and actionlint were clean, and each reported its planted defect.
  • RULE 10: 20-DEADLOCK.md. The one new lock is a leaf, and the abort drain runs where the completed switch's already does.
  • RULE 74: static/STATIC-WITNESS.txt. (*Loop).run and (*Worker).run compile to identical instructions on arm64 and amd64, so the steady path is unchanged. B1 still has to re-pin its B arm to 7aeb2ec, because the pause, resume and switch-abort paths changed.

B1 still holds the merge (body §10). Merge order: this PR, then #696.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Review round 4.2 addressed two minor findings, both on the stall write-up. There is no code change and no new measurement; the head is still 7aeb2ec and main is still 9f4d89b1. The body (TL;DR, §6, §9, §11, §12 and §13) and the round-4 comment were corrected in place; the comment's correction is marked.

1. Who ran beside the stall containers. Round 4 said another lane's container ran beside every one of the 80. That was wrong. None of the 80 ran alone:

  • 53 had another lane's container beside them.
  • 40 had one of this lane's own containers beside them: its -race suites, G4 mutants and CI-interlock cases, run as a second queue beside the stall schedule under the same two-slot lock.
  • 27 had only this lane's own containers beside them.

So part of the load that round 4 blamed on other lanes was this lane's own, and that was not disclosed. The own-lane neighbours fell evenly across arms (8 MiB: A 6/10, B 16/30, C 6/10; 128 MiB: A 6/15, B 6/15), so they do not bias the comparison, and the registered verdict stands.

While checking, the same misattribution turned up twice more, and both are now fixed:

  • The iouringpkg-otherskip interlock case (7 data races, ring ENOMEM at 8 MiB) ran beside two of this lane's own io_uring stall containers at 128 MiB as well as one other lane's container (§9, §13).
  • Seven of the eight round-4 suite runs at c12d1b8 ran beside this lane's own stall containers (§6).

2. Figures no script regenerated, and the TL;DR. Every stall figure outside the registered STALL-TALLY.txt now comes from python3 round4/stall/context.py (writes round4/stall/STALL-CONTEXT.txt) and is labelled as not registered. That covers:

  • the neighbours;
  • the host load;
  • the control stall's ledger: 1,501 detached and 1,501 adopted at R1, answered +19.9 to +20.3 ms after the call;
  • the cross-session comparisons with D2, whose counts are recomputed from D2's raw rounds.

Other changes:

  • The TL;DR now says the control was weak: 0/30 against 1/10, p = 0.25.
  • "The control stalled less today" is withdrawn. 1/10 against D2's 8/30 on the same binary gives a two-sided p of 0.40.
  • One figure was wrong: D2's load range, given as 1.2-3.4, is 1.2-4.0 (the sampler's maximum is 3.97).

Verdict, unchanged. The stall is EXCLUDED at 8 MiB by the registered rule, with an upper bound of 9.5% per container. The positive control stalled only 1 of 10 times, so the result cannot exclude a rate that low.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Review round 4.3 answered three minor findings. None was on product code. There is no code change: the head is still 7aeb2ec and main is still 9f4d89b1. The body's §6, §9, §11, §12 and §13 and its TL;DR were edited in place, and so was #708.

1. #708's main baseline came from the loaded re-gate alone (fixed; #708 edited). #708 said main "failed TestRampH1Async in 3 of 4 containers and TestRampH1Sync in 2 of 4" with 6,615 read: reset. Those are the r6/fix re-gate's rows (2026-09-26). That run had other lanes' native load on the host, and two of its four main containers are P7's main arm, which this PR marks contaminated. It left out the r6 gate (2026-09-20). There, main passed both ramps in 6 of 6 containers (12 of 12 ramp runs), with 1,556 read: reset in 13,885,363 requests. #708 now shows the two runs separately and says to classify a failure against the r6 gate. Every one of main's failing ramps served at most 458,167 requests, against 1.06M-1.28M per ramp in the r6 gate, so #708 now also asks for each ramp's throughput. Per container: round4/r43/RAMP-BASELINE.txt (python3 r43/tools/ramp-baseline.py).

While checking, one slip turned up in round4/30-CORRECTIONS.md §6. It gave "Sync ok=219,754 and 293,984" as P7's main-arm throughput, but 293,984 is the branch container p7-a02. Main's two P7 Sync ramps served 219,754 and 1,056,532 requests. The body only ever cited 219,754, and the contamination marking stands. The load shows in both arms: 175,718-1,056,532 requests per ramp in the re-gate, against 955,898-1,279,432 in the r6 gate.

2. A FAIL on FINAL was decided by an instrument that was not registered (disclosed in §6). 00-PLAN.md §6 said a contemporaneous base re-run decides any FAIL on FINAL that base does not show. For TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen/drain_stops_while_held (s2, 8 MiB, ring ENOMEM at engine start), no base re-run was made. Real CI and the ENOMEM diagnosis decided it instead, and round 4 did not call that a deviation. Checking this turned up a second departure from the same section, which the review did not raise. The plan says to stop before pushing on an attributable FAIL, but 7aeb2ec was pushed at 03:07:33Z and its suites only started at 03:09:15Z, so that stop could never have held the push. The conclusion is unchanged:

  • the test's file is byte-identical to main's;
  • base passed the subtest 4 of 4 in the re-gate;
  • CI passed it at 8 MiB (below).

No late base re-run was made. On this host, a local 8 MiB run is starved of ring memory by the other root containers (the charge is per UID), and a controlled run would need a ring hog that starves the other lanes. Details: round4/r43/10-DEVIATION-AND-CORRECTIONS.md.

3. The strict CI gates had one real-CI sample (re-run three times; recorded in §9). Run 36290508684 was re-run three times on the same bytes (attempts 2-4, merge commit 85fb140 each time), and all four attempts were 9 of 9 green. In every attempt:

  • the unit job had PASS 435, FAIL 0, SKIP 5 (the allow-list), and the adaptive job had 111 / 0 / 0;
  • the 9 tally lines were identical, and no job logged a DATA RACE;
  • the engine/iouring step took 155.7-156.7 s of its 300 s;
  • drain_stops_while_held passed twice at 8 MiB;
  • TestLingerTransplantConserves had err=0;
  • the default-linger storm passed. It logged 1, 0, 1 and 0 dial errors, which it does not assert on.

Zero failures in four runs only puts a flake rate below about 53% per run, so this corroborates and does not prove. Census: round4/r43/CI-RERUN-CENSUS.txt.

The merge order is unchanged: this PR first, then #696. Until #696 lands, the engine/iouring step can go red on celeris#691 at main's own rate. Main runs the same test in its root step today, and the #657 and #662 steps after it run anyway.

@FumingPower3925

FumingPower3925 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Change of instrument for B1: maintainer decision, 2026-09-27 ~08:45Z.

B1 was the pre-registered laptop timing check (S1b run2). It never took a timing sample. Its quiet-host gate never opened: 585 polls from 06:17 to 08:44Z, 1-minute load median 3.3 and never below 1.5, because of interactive use of the host. B1 ended not_run, and it says nothing either way about a regression.

Replacement: a same-run A/B/A′ benchmark on the bare-metal cluster, both arches. Run 36308789830.

  • A = main 9f4d89b; B = this PR at 7aeb2ec; A′ = main again, as an in-run twin. Each engine runs A, B, A′ back to back, so a linear drift cancels, and A′/A is the noise floor for each arch and cell.

  • Cells: 195 per arch. They cover the H1 scenarios, including churn and keep-alive, plus the WS/SSE loop cells, for epoll, io_uring and adaptive.

  • The read rule was pre-registered before any data (PREREG sha256 76529841…). It uses the per-cell B/A delta against the A′/A floor, jackknife and bootstrap support, RPS and CPU/req, per arch. Outcomes:

    • NO REGRESSION: the PR can merge.
    • REGRESSION at named cells: the merge is blocked.
    • MARGINAL: one pre-registered replication of those cells.

    Correction (09:25Z): the floors are coarser than first stated here. Only amd64 RPS reaches the ~1.5% figure; its engine floor is 1.55%. The other engine floors in this design are amd64 CPU/req 3.46%, arm64 RPS 3.82% and arm64 CPU/req 2.40% (PREREG §7, design_floors.out row XS). Per-cell class floors run from 1.5% to 16.6%. In a smoke test, an injected -5% engine shift was only MARGINAL on arm64. So this run can exclude a regression of a few percent, not B1's ~1%. The expected cost, ~0.02%, is far below either bar.

  • Expected, from static analysis: no regression. This PR adds 5 instructions per idle event-loop iteration and none per connection, about 0.01-0.02% per operation.

B1's witnesses still stand, as supporting evidence (not as data for this run):

  • E2 static diff: 0 unexplained differences.
  • E1 socket witness: 8/8 HOLD.
  • E3b syscall census: structural syscalls per op are identical between B and A.

The verdict will be posted here when the run completes (ETA ~12:15Z).

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Cluster A/B verdict: NO REGRESSION. Run 36308789830 completed successfully at 12:11Z. It was read with the pre-registered ab674_verdict.py; the PREREG and its tools were re-hashed OK before the read.

Remaining review minors (the ci.yml ramp-quarantine comment still says main fails the ramps, contradicting the corrected #708; the stall write-up wording) are tracked in a v1.6.0 follow-up under the maintainer's two-round review cap. Merging.

@FumingPower3925
FumingPower3925 merged commit 49d2726 into main Sep 27, 2026
40 checks passed
@FumingPower3925
FumingPower3925 deleted the fix/celeris-662-defer branch September 27, 2026 12:20
FumingPower3925 added a commit that referenced this pull request Sep 27, 2026
… hijack accept test

#674 (49d2726) added a deferAccept parameter to createListenSocket and
updated every caller then on main, passing true to keep the old behaviour
(TCP_DEFER_ACCEPT was always set). This PR's
TestAcceptOfANumberAHijackReleasedIsOrderedAfterTheHijack was written
against the old signature, so after merging main the engine/epoll test
package no longer compiled (GOOS=linux go vet: "not enough arguments in
call to createListenSocket"). Pass true, as #674 did for its siblings.

Checked: GOOS=linux go build ./... and go vet ./... on amd64 and arm64,
rc=0 (both failed on the vet before this change).
This was referenced Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment