fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) - #674
Conversation
Not merging: two of three reviews returned request_changes, and I re-verified the majors at
|
Round 3 status: not pushed, because its own regenerated evidence does not clear the bar it setA host reboot (an OS upgrade) wiped round 3's measurement logs mid-run. The code survived as two Where no switch is reachable (memlock 8 MiB): the gated engine keeps Where a switch IS reachable (memlock 128 MiB — the typical host, running the default engine): So B1 is narrowed, not answered: the default engine still carries a measured churn cost Also caught before pushing: round 3's new tally in the adaptive CI job ended its regex in Withdrawn from the current body: the error-census table (its recorder saturated at 200 Next: before another code round, checking whether the trade-off can be dissolved instead of |
…se drain (celeris#662) #663 taught PauseAccept to serve what its accept queues hold, which fixed the queued case. It cannot reach a connection whose handshake completed but which has sent no data: createListenSocket sets TCP_DEFER_ACCEPT=1, so the kernel holds that connection as a TCP_NEW_SYN_RECV request socket and keeps it out of the accept queue entirely. The drain never sees it, the listen-socket close orphans it, and the client's first request is answered with a reset. Three failing-first tests, in both engines and on the adaptive switch: - TestPauseAcceptKeepsHandshakedIdleConnections dials 8 connections that write nothing, asserts the accept count BEFORE the pause -- 0 on main, which is the defect -- pauses, and only then sends the requests. RESET 8 of 8 on main, on both engines and at both memlock shapes. - TestListenSocketDoesNotDeferAccept asserts the property on the socket itself. It calls createListenSocket directly rather than reading a loop-owned listenFD from the test goroutine, which -race would flag as a fault in the test instead of the engine. - TestSwitchKeepsHandshakedIdleConnections promotes epoll -> io_uring with 16 such connections established: RESET 16 of 16. It is guarded by requireUpSwitch, so the adaptive CI job's CELERIS_REQUIRE_UPSWITCH=1 turns its skip into a failure. TestPauseAcceptIdleControl is the same rig without the pause. It passes on main, which attributes the paused arm's failure to the pause rather than to the rig's unusual write-nothing clients. Each engine rig also logs the kernel's own witness, TcpExtTCPDeferAcceptDrop, whose delta over the dials is 8 while the option is set and 0 once it is gone. The 1 s deferral timer is these rigs' only clock, and it is a vacuity hazard rather than a flake hazard: past about a second the kernel's SYN-ACK retransmit creates the children itself and main would pass for the wrong reason. The measured phase is bounded two orders of magnitude below that, and the elapsed time is asserted, so an overrun fails the test loudly instead of hiding inside it. The setsockopt is deliberately left in place. This commit only proves the loss class.
…eeps the clients that had not yet sent a request (celeris#662) the other half: with TCP_DEFER_ACCEPT set, a connection whose handshake has completed but whose client has sent nothing is never IN that queue. The kernel keeps it as a TCP_NEW_SYN_RECV request socket and accept4 answers EAGAIN, so acceptQueuedOnPause cannot see it, the listen-socket close orphans it, and the client's first request is met with a reset. The engine counts no accept, no close and no error. Every adaptive switch pauses the outgoing engine, which is where this bites: celeris#660's CI job saw a 2048-connection ramp lose about a third of its connections across a promotion, every sampled error a reset. The option is not removed outright. Pricing it first (30 rounds per arm, two memlock shapes, counterbalanced, each cell against its own A/A floor) showed dropping it costs +14-21% ns/op and +2.4-3.5 us of server CPU per connection on ordinary churn, +67-200% on connect-and-never-send, and +1974% ns/op on io_uring at the CI's one-worker memlock. Keep-alive is flat in both units (p=0.94 and 0.20), so the cost is confined to the accept path. 19 of 24 co-primary cells regress at Holm-adjusted p<1e-8 with bootstrap CI lower bounds far above the pre-registered +2% materiality bar, against A/A floors that do not move. So the option becomes configuration, defaulting to ON: resource.Config.DisableDeferAccept, mirrored on celeris.Config, threaded into both engines' createListenSocket. adaptive.New sets it on BOTH sub-engines, because adaptive pauses on every switch and so never gets to keep an option that hides connections from a pause. An engine that never pauses keeps the throughput. A standalone epoll or io_uring engine whose owner calls PauseAccept must set the flag; that residual exposure is called out in the field doc and on the issue. Tests: the two engine rigs build their engine the way one that intends to pause must be built, and TestListenSocketDeferAcceptFollowsConfig pins both directions of the config on the socket itself -- including that the default still sets the option, so the measured throughput default cannot be flipped silently. The adaptive rig needs no config of its own: if it fails, adaptive stopped setting the flag. CI: TestRampH1Sync and TestRampH1Async come off the `adaptive` job's -skip list. `-skip` is a regex, so a rename would drop them silently and go test would still exit 0; the job now tallies their PASS lines and fails if it does not see exactly two, the same interlock the `iouring` job uses. Prose corrected where it described the old unconditional behaviour: both acceptQueuedOnPause doc comments now say the drain can only reach the accept queue, epoll's "Not covered" list distinguishes what is inherent to closing a listen socket from what is configuration, and the nodelay-inherit test says which socket asked for the deferral.
…not their only cause This PR's first commit lifted the celeris#662 quarantine on TestRampH1Sync and TestRampH1Async. CI disagreed, twice, so the lift is reverted here. Two runs on GitHub-hosted runners with the fix applied and memlock unlimited (run 35082406457, first attempt and a re-run of identical bytes): attempt 1: --- FAIL: TestRampH1Sync (65.02s) --- FAIL: TestRampH1Async (56.96s) attempt 2: --- FAIL: TestRampH1Sync (57.07s) --- FAIL: TestRampH1Async (56.86s) 91 PASS / 2 FAIL / 0 SKIP both times, and those two were the only failures in the package. Reproducible, so not the celeris#670/#673 flake situation -- which is a different pair of tests in any case. celeris#662's own loss class IS fixed on that runner. The issue recorded sync phase 0 holding 1363 of 2048 conns with 685 errors, every sampled one a "read: ... connection reset by peer". With this fix phase 0 holds 2048 of 2048 with err=0 in both variants, and the job-wide count of read-resets is 0. The 52 resets that remain are at DIAL -- a handshake still in progress when the listen socket closes, which acceptQueuedOnPause documents as not covered and inherent to closing a listen socket. What still fails is later phases that only these two tests exercise: * async phase 4: io_uring=1316 of 2048 after 20s, err=0 across 938,818 requests -- nothing is dropped, the transplant just does not finish (#657) * sync phases 3 and 5: connections stranded on the down-switch Neither is accept-time loss, so both names go back on -skip under celeris#657, and the PASS tally interlock is removed with them: it demanded two PASS lines that will not appear. The job's comment now records this measurement so the next attempt starts from it rather than repeating it. Leaving them un-skipped was the alternative and it is worse: the Adaptive job would be permanently red, and a job that can never be green stops being read -- which is how this coverage was lost in the first place.
…ally pause (celeris#662) Adaptive is the DEFAULT engine on Linux (resource.defaultEngine), so setting DisableDeferAccept on every adaptive engine put the option's measured churn cost on the default configuration -- including the many adaptive engines that can never switch, and so never pause, and so can never hit celeris#662. The flag is now gated on switchPossible(): the conns-per-worker up-switch being reachable (epoll start + io_uring viable + non-h2c), OR an io_uring start, whose always-on error-rate revert can pause it too. An old kernel, a memlock below one io_uring worker's rings, or Protocol H2C keeps the option and pays nothing. connSwitchEnabled is now computed from the same predicate so the gate and the controller cannot drift apart, and the override is logged rather than silent. An engine that CAN switch still pays: the pause is what loses the connection, so only an engine that never pauses can safely keep the option. Also in this commit: * The engine rigs ship BOTH arms. The deferred arm asserts the mechanism the fix rests on -- the kernel holds a handshake-complete connection that has sent no data out of the accept queue -- so the two arms make opposite assertions about the same counter and neither can pass if the flag stops reaching createListenSocket. The doc comments no longer claim these files pass on main: they name a field main does not have and cannot compile there. adaptive's switch test names no new API, and is the artifact that fails on unmodified main. * CI can no longer go green by skipping them. The unit job gets a PASS-count interlock over the eight celeris#662 engine tests (-v, memlock raised, CELERIS_REQUIRE_IOURING_WORKERS=1), copying the iouring job's interlock from celeris#664; the adaptive job gets one over the three gate tests. io_uring's two unguarded skips now route through skipOrFail656, and the adaptive standby skip -- which sat after requireUpSwitch and bypassed CELERIS_REQUIRE_UPSWITCH entirely -- honours it. * The ramp error census no longer saturates. recordErr guarded on map SIZE while its keys embedded the client's ephemeral port, so it stopped counting after 200 connections, repeats of existing keys included. Counting is now unbounded and per class; only the verbatim exemplars are capped. Every by-kind figure that rested on the old map is withdrawn rather than restated. * Public documentation is made true: Server.PauseAccept, AcceptController and the README said existing connections continue to be served, which is not so for a handshake-complete connection that has sent nothing. The residual for standalone engines is celeris#675. * io_uring's acceptQueuedOnPause doc lists both inherent residual states epoll lists, not one; resource.Config documents the lost connect-and-never-send shield as a resource question, not only a throughput one. * The A/B benchmark that priced the option is committed as test/deferab, with adaptive arms added, so the measurement that chose the design can be re-run from the repository. Refs #662 Refs #675
…the option switchPossible() already documented erring toward correctness on the CELERIS_ADAPTIVE_START=iouring escape hatch. It errs the other way in exactly one place, and leaving that to be discovered is the habit this review was about: ForceSwitch is exported and calls performSwitch directly, bypassing the controller the predicate reasons about, so calling it on an engine the gate judged non-switchable does pause a sub-engine and can drop a handshake-complete connection that has sent nothing. No production path reaches ForceSwitch (it is documented "for testing"). Widening the gate to cover it would put the measured churn cost back on every non-switchable adaptive engine -- the default configuration the gate exists to protect -- to insure a method nothing in production calls.
The gate turned TCP_DEFER_ACCEPT off on both sub-engines whenever a switch was reachable, which put the option's churn cost on the default engine of every switch-capable host. The fix now lives in the engines' pause instead (clear the option on the pausing listener and keep it open for a linger), so adaptive no longer needs to decide anything about the option. adaptive/engine.go is back to main's bytes: switchPossible, upSwitchPossible and the DisableDeferAccept override are gone, and connSwitchEnabled is computed exactly as main computes it. The caller's DisableDeferAccept now reaches both sub-engines untouched. The gate's test file and its two names in the adaptive job's tally go with it; the tally itself is rewritten with the CI interlocks later in this series.
…pt pause (celeris#662) A listener with TCP_DEFER_ACCEPT holds a handshake-complete connection that has sent nothing outside the accept queue until the kernel's SYN-ACK timer promotes it, about one second after its SYN. A pause that closes the listener before then resets it. The engines will instead clear the option on the pausing listener, keep it open for a linger measured from that clear, and only then run their drain-and-close. This package holds what the epoll and io_uring engines share for that: - Linger (1500 ms) and three test hooks, ClearOnPause, GuardEnabled and ObserveDelay, all atomics read only in the pause step; - SynackRetriesZero and GuardWanted: TCP_SYNCNT=1 is applied to a pausing listener only when net.ipv4.tcp_synack_retries was read successfully and is 0, because the kernel rejects TCP_SYNCNT below 1 and a guard applied on an unreadable value could never be undone; - Clear, Restore and GuardSyncnt, and Enter/Leave, the ACTIVE->LINGERING and LINGERING->ACTIVE transitions, with the deadline taken after the setsockopt returns; - PauseState, which BeginPauseAccept fills off the loop threads (the sysctl is read once per pause, there); - process-wide counters for lingers, closes, aborts, setsockopt failures, guards and unreadable sysctls. It lives under the module's internal/ rather than engine/internal/ because the adaptive package's tests need its hooks and counters, and Go only lets code under engine/ import engine/internal/.
…s the listener (celeris#662, celeris#675) A pause used to close each loop's listener as soon as the loop saw it. With TCP_DEFER_ACCEPT on, which stays the default, a client whose handshake had completed but which had not sent its request yet was still a request socket outside the accept queue, so the close reset it, and the engine counted nothing. Now each loop, on its own thread, clears the option on its own listener, takes its deadline deferlinger.Linger (1.5 s) after that clear, and keeps accepting and serving through the normal path until the deadline. The kernel promotes a deferred connection about one second after its SYN, which raises an edge on the still-registered listener like any other accept, and anything that arrives after the clear enters the accept queue at once. At the deadline the loop runs the unchanged drain-and-close, now in closeListenerAfterDrain. A resume during the linger restores the option on the same descriptor. A listener built with DisableDeferAccept, or a zero linger, closes at once as before. - BeginPauseAccept starts a pause and returns; PauseAccept is Begin plus a wait that ends when every listener is closed, when a resume withdraws the pause, when every loop has exited, or after Linger + 1 s. - listenFDClosed is true exactly while paused with no listener, and loop shutdown now sets it, so a PauseAccept racing Shutdown returns. - The epoll_wait timeout is capped at the time left to the deadline. While a loop listens its wait is already a few ms, so this is defence.
…ses the listener (celeris#662, celeris#675) The io_uring half of the epoll change. At a pause each worker, on its own thread, clears TCP_DEFER_ACCEPT on its own listener, takes its deadline deferlinger.Linger (1.5 s) after that clear, and keeps its multishot accept armed and keeps serving until the deadline; then the unchanged cancel/drain/close runs, now in closeListenerAfterDrain. A resume during the linger restores the option on the same descriptor, whose accept was never cancelled. A listener built with DisableDeferAccept, or a zero linger, closes at once as before. - BeginPauseAccept starts a pause and returns. It promises no prompt start: a plain HTTP/1 worker has no eventfd poll armed and sees the flag at its next ring wakeup, up to 100 ms later. Nothing depends on that, because each listener's deadline is taken at its own clear. - PauseAccept is Begin plus a wait that ends when every listener is closed, when a resume withdraws the pause, when every worker has exited, or after Linger + 1 s. - rearmAcceptIfPending gates on the listener only. It used to refuse while paused, which would leave a lingering worker whose accept re-arm hit a full SQ ring deaf until its close. TestAcceptRearmRetriedAfterSQFull is rewritten to that rule (re-arm while the listener is open, lingering included; never once listenFD < 0), and TestAcceptRearmedDuringLinger drives the retry inside a real linger entered through the pause step. - adaptiveTimeout is capped at the time left to the linger deadline; an idle worker would otherwise overshoot it by up to 100 ms. - listenFDClosed is true exactly while paused with no listener, and worker shutdown now sets it, so a PauseAccept racing Shutdown returns.
…for its linger (celeris#662) A switch used to call PauseAccept on the engine it left and wait for its listeners to close. That pause now lingers for about 1.5 s with TCP_DEFER_ACCEPT cleared, so the outgoing engine can serve the clients that completed their handshake on it before the switch but had not sent their request yet. Waiting for that inside performSwitch would hold e.mu, and with it Metrics() and the next switch, for the whole linger. Both pause sites in performSwitch (the abort path for a freshly built standby, and the final pause of the old active) now call beginPause, which uses the sub-engines' BeginPauseAccept and falls back to PauseAccept for test fakes. It is still issued under e.mu, so a later switch's ResumeAccept of the same engine is ordered after it, and a resume during the linger restores the option on the same listeners. Comment fixes: "PauseAccept itself caps its wait to 2s" was wrong, and the freezeState release rationale no longer depends on a pause that waits.
…ts instant opt-out (celeris#662, celeris#675) The public docs said a client that had connected but not yet sent its request was not carried across a pause, and told callers to set DisableDeferAccept. That is no longer true: the pause clears TCP_DEFER_ACCEPT on each listener and keeps it open for about 1.5 s first. - Server.PauseAccept, engine.AcceptController.PauseAccept and README: the pause carries such clients, blocks about 1.5 s and keeps admitting meanwhile; what it can still reset (a handshake in flight at the close, a lost or unanswered SYN-ACK retransmit); DisableDeferAccept for a pause that closes at once. - Config.DisableDeferAccept and resource.Config.DisableDeferAccept: the mechanism, the linger, and the opt-out. No cost figures in code comments. The claim that a never-sending connection "never becomes a socket the engine owns" was wrong: the kernel promotes it about one second after its SYN, and the engine accepts it then. - acceptQueuedOnPause (both engines): the drain cannot see a deferred connection, which is why the pause lingers before it; the remaining residuals. - createListenSocket (both engines), and the engines' BeginPauseAccept / PauseAccept docs without internal package names. The io_uring stepAcceptPause/closeListenerAfterDrain block moves above acceptQueuedOnPause's doc comment, which it had split.
…that must lose (celeris#662, celeris#675) pause_accept_linger_linux_test.go, the same tests in epoll and io_uring, all on the default configuration (TCP_DEFER_ACCEPT on), none with subtests: - T1 TestPauseAcceptServesDeferredIdleConnections: eight clients connect and send nothing, the pause starts inside the kernel's one-second timer, and they write only after every listener has closed; all eight are accepted during the linger and served, the pause lasts at least a second, every listener reads the option off during the linger, and a later dial is refused. - T2 TestPauseAcceptServesArrivalsDuringLinger: a silent dial every 15 ms through the pause; none that connected before the close's last 60 ms is lost, and every one after the clear (seen on the sockets) is accepted within 100 ms. - T3 TestPauseAcceptLingerAnchoredAtClear: the loops act on the pause 800 ms late (deferlinger.ObserveDelay); still nothing lost, so the deadline is measured from each listener's clear. - T4 ...SynackRetriesZero: T1 at net.ipv4.tcp_synack_retries=0, where the pausing listeners read TCP_SYNCNT 1. Skips elsewhere; CELERIS_REQUIRE_SYNACK0=1 makes the skip a failure. - T5 TestResumeDuringLingerRestoresDeferAccept: same SO_COOKIEs across a resume during the linger, the option back on, a fresh silent client hidden again, and PauseAccept returns at once. - T6 TestShutdownDuringLingerIsPrompt: a shutdown during the linger releases PauseAccept's wait at once. - N1 linger 0 -> RESET, N2 clear disabled -> arrivals lost, N3 guard disabled at synack 0 -> RESET: each passes only by observing the loss. Every rescue test checks its premises first: hidden within 200 ms and TCPDeferAcceptDrop moved, dials well inside the one-second timer, and tcp_migrate_req 0 (CELERIS_REQUIRE_LINGER_PREMISES=1 turns that skip into a failure). Existing tests: TestPauseAcceptDeferredIdleConnectionsAreHidden becomes the pause-free premise control; TestPauseAcceptKeepsHandshakedIdleConnections stays as the DisableDeferAccept (instant, lossless) arm; the queued rig waits Linger + 2 s for the close. Cost figures are gone from the test headers.
…(celeris#662) switch_idle_conns_linux_test.go names no API main lacks, so it can be copied onto main or onto the round-3 gate: - TestSwitchKeepsHandshakedIdleConnections is rewritten for the new truth: the silent clients are HIDDEN before the switch (TCP_DEFER_ACCEPT on), the previous warm-up linger is waited out first, and the clients stay silent until every listener the outgoing engine had has closed, then write; all sixteen are served. It fails on main, where the outgoing listeners close at once. TestSwitchKeepsHandshakedIdleConnectionsOnRevert is its io_uring->epoll twin. - TestAdaptiveListenersKeepDeferAcceptAcrossSwitches: socket witness (/proc/self/fd, SO_COOKIE, TCP_DEFER_ACCEPT) before any switch, on the incoming engine after a promotion, on the same epoll listeners after a flap back inside their linger, and after full cycles. The round-3 gate fails it. These no longer need the controller's up-switch (ForceSwitch performs the switch), so they run at the 8 MiB memlock shape too, with one io_uring worker; the controller is frozen so only the test switches. linger_switch_linux_test.go: - TestPerformSwitchDoesNotBlockOnLinger: ForceSwitch under 100 ms and Metrics() under 50 ms while the outgoing listeners are still open. - TestFlapWithinLingerKeepsListeners: a switch back inside the linger resumes the same listeners (same SO_COOKIEs) with the option back on. - TestLingerTransplantConserves: keep-alives transplanted across a promotion while churn keeps landing on the lingering epoll; detached == adopted, no silent-loss bucket moves, accepts == closes == hooks. Repairs: TestAdaptiveSwitchVsAcceptChurn sets the linger to 0 and waits for the outgoing close after each switch, and asserts from deferlinger's close counter that listen sockets did close under the accepts (a storm of non-blocking switches otherwise never closes one). flapScenario spaces its flaps Linger + 0.7 s apart so each flap closes and re-creates listeners as before.
…from the comments (celeris#662) Every benchmark now reports, read through /proc/self/fd before the timed region, how many LISTEN sockets the engine has on its port (listeners) and how many of them read TCP_DEFER_ACCEPT off (listeners_defer_off). With the gate deleted, the steady state must have the option on every listener, on every engine including adaptive at both memlock shapes; this makes that a per-round reading next to deferdrop/op instead of an inference. The package comment and the adaptive arm's comments described the round-3 gate, which no longer exists.
…d run the synack=0 guard (celeris#662) unit job: - The celeris#662 step now runs the linger tests (T1-T3, T5, T6), their negative controls (N1, N2), T7 and its unit twin, and the DisableDeferAccept arm with its premise and socket checks, -v with skipping forbidden. It asserts that exactly the listed names ran and passed at top level (`=== RUN Name` with nothing after it, and `--- PASS: Name (`), and fails on any `--- SKIP` line at any indentation, because a skipped subtest still lets its parent report PASS. The expected count is derived from the name lists. - New step: net.ipv4.tcp_synack_retries=0 (restored on exit), running T4 and N3 in both engines with the same tally and skipping forbidden. adaptive job: the tally regex ended in `)$` and could never match `--- PASS: Name (0.01s)`, so the job could never go green. It now counts the seven celeris#662 adaptive tests the same way, and a SKIP line for any of them fails the step. Every step sets pipefail explicitly as well as through `shell: bash`.
…red the option (celeris#662) TestAdaptiveListenersKeepDeferAcceptAcrossSwitches flapped back to epoll microseconds after the promotion, before any epoll loop had acted on the pause, so the flap withdrew a pause no loop had seen: nothing was cleared, nothing needed restoring, and the mutation record showed the test passing with the restore removed. It now waits until every epoll listener still open reads TCP_DEFER_ACCEPT off (the clear) before it flaps, and logs that state, so the flap exercises the restore on the same listeners.
…ines (celeris#662, celeris#657) The pause keeps a listener open for about 1.5 s with TCP_DEFER_ACCEPT cleared, so for that long the engine a switch is leaving keeps its SO_REUSEPORT share of NEW connections. Nothing pinned where those go, and that is exactly what failed this PR's base-vs-branch suite gate on 2026-09-18: with the linger in and celeris#657 open, an async 2048-connection ramp left 972-1170 conns on the outgoing epoll after 20 s, err=0. Not a loss -- placement. celeris#657 is now fixed (#681, #687) and this is the test that holds the joint property, so the interaction cannot regress silently. TestLingerArrivalsReachTheIncomingEngine, in engine/epoll and engine/iouring, with an async handler and the default configuration: * BeginPauseAccept, then wait for the socket witness to read the option off every listener -- the linger has begun; * PHASE A: six connections arrive, are served, and go silent while NO drain is set. Their dispatch goroutines park with nothing owed for them (epoll: askAtPark returns early with no drain; io_uring: runAsyncHandler's claim is not reached). This is the flap case, a second switch inside the first one's linger; * StartTransplant, still inside the linger; * PHASE B: six more arrive, are served, and go silent with the drain set. They join the live set at the TAIL, past the sweep cursor -- the case the cycle rule exists for (celeris#657 R2 MAJOR-1); * then every listener closes and the engine must hold none of them. It asserts what a client and an operator can each check: no dial refused and every arrival accepted by THIS engine (OnConnect is the witness), every arrival answered 200, every one adopted by the target, ActiveConnections back to 0, and the hand-off ledger balanced (io_uring also requires TransplantHandoffInFlight 0). The target serves what it adopts. A target that only holds the descriptor cannot tell a connection that was MOVED from one that was DROPPED, because either way its client never gets an answer -- the first draft of this test read 12 timeouts and 12 adoptions and could not say which had happened. Both controls lose, measured in one container each, mutation restored and sha256-verified (evidence celeris-662/r6/rebase/logs/c05..c08): * linger 0 (celeris#662 off): 12 of 12 dials refused, 0 accepted, 0 served, on both engines. * the sweep and the park-boundary ask removed (celeris#657 off): epoll adopts 0 of 12 and still holds 12; io_uring adopts 6 of 12 and still holds the 6 of phase A. io_uring's phase B moves without the sweep because its own park claims it, which is why phase A is in the test: it is the only shape on that engine that the sweep alone carries. CI: the name joins the celeris#662 linger step's tally, which runs both packages with -v, forbids every skip and asserts the exact PASS count.
…useAccept instead of polling, and log each pause warning once per engine (celeris#662) Review round 4 of PR #674, four product findings in the pause linger: * The linger deadline was a wall-clock reading (time.Now().UnixNano()), while PauseAccept's own bound is monotonic. A backward step of the wall clock during a linger kept a paused listener open and accepting after PauseAccept had returned; a forward step ended the linger early and brought the resets back. deferlinger now has a clock of its own, Now: nanoseconds since the package's initialisation, read with time.Since, so on the monotonic clock. Every linger deadline is taken and read through it, as a distinct Deadline type, so a deadline cannot be compared with a wall-clock reading without a conversion. TestEnterDeadlineIsOnTheMonotonicClock fails on 1d90b5d (its deadline is in the Unix-nanosecond domain); TestClockIsMonotonic pins that the epoch carries a monotonic reading. * A standalone PauseAccept slept 1 ms at a time through the whole 1.5 s linger: about 1,500 wakeups per pause, 65-73 ms of CPU on an idle io_uring engine. It now waits on a per-engine broadcast channel (PauseState.Changed/Notify): a loop's or worker's pause close and its exit set listenFDClosed and then notify, and ResumeAccept notifies. The channel is taken before the flags are read, so no wakeup is lost. A 100 ms re-check stays only for the one change that is not notified, a pause that lands on a listener a resume has not re-created yet. The wait mutex is a leaf lock. TestPauseAcceptWaitIsWoken (both engines) pins that the wait ends on a notification, after a close and after a resume. The run loops' bodies are unchanged. * Begin warned on every pause when tcp_synack_retries was unreadable, under the adaptive engine's lock on every switch, and a failed clear, guard or restore warned per listener per pause, although the comments said "logged once". Each is now logged once per engine and still counted every time. TestPauseWarningsAreLoggedOncePerEngine fails on 1d90b5d. * lingerTimeoutMs(-1) returned -1, so a lingering epoll loop handed an infinite wait would have slept through its deadline. It is capped too. TestLingerTimeoutNeverBlocksPastTheDeadline fails on 1d90b5d.
…itch lingers it (celeris#662) A switch that builds the lazy standby and then finds a driver registered during the build aborts and pauses the fresh engine. That pause lingers for about 1.5 s with the engine in the SO_REUSEPORT group, so it accepts its share of new connections -- about half on a four-plus-four split -- and nothing moved them: they stayed on a standby that no switch may ever make active (review round 4 of PR #674). The abort path now gives the fresh engine the drain a completed switch gives the engine it leaves, toward the engine that stays active. The linger itself stays: closing at once would reset a connection whose handshake the fresh engine completed during the build, the celeris#662 class. The call runs under e.mu exactly where the completed switch's applyTransplant already does. TestSwitchAbortedByADriverDrainsTheFreshStandby fails on 1d90b5d (the io_uring standby keeps the connections it accepted during its linger) and passes here: 17 of 32 landed on it, 17 detached, 17 adopted by epoll.
…etry their io_uring start on ring ENOMEM, and log the accept margin (celeris#662) Review round 4 of PR #674, test findings: * runLingerArrivalsL662 (both engines) stopped its dialer and its /proc/self/fd sampler only on the success path. A t.Fatal before that left the dialer dialing a freed port, one every 15 ms, for the rest of the package run, where a later test's engine could reuse it. Both are now stopped by a sync.OnceFunc that the normal path and t.Cleanup call. * The io_uring linger rigs (startLingerL662, startLingerAsyncL674, runPauseIdle662) sent an engine start that failed on ring ENOMEM straight to skipOrFail656. They now go through startRingRetried662, which retries such a start for as long as retryRingENOMEM does, as startTestEngine and the two celeris#639 tests already did. * ioUringUsableHere cached its first answer in a sync.Once, so a probe that met a transient ENOMEM turned every later budget failure in the process into a skip. A ring ENOMEM now means io_uring works and the budget is held, which is what skipUnlessIOUringUnusable exists to fail on; only a positive answer is cached. * T2 logs its largest dial-to-accept time next to the 100 ms bound, so its margin on a CI runner is visible.
…, and storm the switch with the default linger too (celeris#662) Review round 4 of PR #674: * TestLingerTransplantConserves logged keep-alive and churn errors but asserted neither, so its conservation check had never been seen to fail. It now fails on any keep-alive error, on any churn failure that is not the reset closing a listen socket inherently allows (RESET, or at the dial), and on more of those than two per outgoing listener. Over 66 earlier runs keep-alive errors were 0 every time and churn resets 0-4. Its churn and keep-alive drivers are stopped by t.Cleanup after an early failure. * TestAdaptiveSwitchVsAcceptChurn sets the linger to 0, so the storm no longer covered the configuration that ships. TestAdaptiveSwitchVsAcceptChurnDefaultLinger storms it with the default linger, at intervals that land both inside a linger (a resume racing the accepts in flight) and after its close (a re-created listener), and requires the fd count to settle, the port to hold exactly the active engine's listeners with the option on, the engine to serve, and both paths to have run (deferlinger Aborts and Closes).
…teps after a red step, and tally the new tests (celeris#662) * The adaptive job's quarantine of TestRampH1Sync and TestRampH1Async named celeris#662 and #657; #657 is closed and this PR closes #662, and the workflow's rule is one open issue per quarantine. celeris#708 now owns it: what the lift takes (>= 6 GitHub-hosted runs of identical bytes) and the evidence so far. The comment no longer carries this branch's history. * The celeris#657 witness step, the celeris#662 linger step and the synack=0 step run `if: ${{ !cancelled() }}`, so a red engine/iouring step (celeris#691, pre-existing on main until PR #696) cannot hide whether they pass. * The linger step's tally adds TestPauseAcceptWaitIsWoken (both engines); the adaptive tally adds TestAdaptiveSwitchVsAcceptChurnDefaultLinger and TestSwitchAbortedByADriverDrainsTheFreshStandby.
… one that can race its re-check (celeris#662) TestPauseAcceptWaitIsWoken as first written withdrew its second pause 300 ms after starting it, exactly three periods of the wait's 100 ms re-check timer. In the round-4 suite at 8 MiB (./engine/iouring, one worker) the timer fired first, the wait returned on it, and the test failed with "the resume did not wake the waiting PauseAccept" although the resume had notified. A single pause's wake source is a race whenever its close or resume lands on a re-check, so the test now asserts the notification over several pauses, where only its total absence fails: * one pause at the default linger: it takes the linger, closes every listener and re-checks no more often than its length allows; * five pauses at a 50 ms linger: the listeners' close wakes the wait; * five pauses withdrawn at 130-330 ms, never a multiple of the re-check: the resume wakes the wait. With notifications removed every one of them ends on the timer, so the mutant that disables Notify is still caught every time. TestNotifyWakesEveryWaiter now checks the notified channels directly: its waiters used WaitChanged, which returns true on the 100 ms re-check too, so it passed without a notification. One message in TestPauseWarningsAreLoggedOncePerEngine said "logged nothing" where the count was only wrong.
…eleris#662) startTestEngine's ENOMEM retry (9a01f71) started each attempt's Listen in a goroutine that closed the shared `done` variable after sending Listen's result. When an attempt failed and the next one began, nothing ordered that goroutine's read of `done` before the next attempt's write, and -race reported it. Round 4's iouringpkg-otherskip interlock case (the engine/iouring step run as CI runs it, at 8 MiB, with another lane's container running beside it) retried starts and got seven DATA RACE reports, each failing the test that made the start. The attempt now captures its own engine and channel.
…M (celeris#662) The other retry sites (the two celeris#639 tests and the linger rigs' startRingRetried662) say when a start was retried; startTestEngine did not, so a run could not show whether its retry path, where the race fixed in the previous commit lived, had been exercised at all.
Review round 4: every finding, what was done, and the proofHead The two majorsReview 1 (major): the keep-alive stall at a one-worker io_uring revert, re-measured on this head. I re-ran D2's own rig and schedule (5k conn/s of churn and 32 keep-alives, four forced switches per container). The runs used one container per observation, counterbalanced, pre-registered and hashed before the first container (
Under the pre-registered rule, the stall is EXCLUDED at 8 MiB, but the control was weak. The rule asks only that the control stall at least once, and it did. This head did not stall in 30 containers. The one-sided 95% upper bound is 9.5% per container, 4.9% per revert. At 128 MiB the bound is 18.1% per container. The only registered comparison is this head against today's control, 0/30 against 1/10, and it is not significant on its own (p = 0.25). Not registered, across sessions ( Corrected in review round 4.2: this comment first said that another lane's container ran beside every one of the 80 containers. That was wrong. None of the 80 ran alone:
The own-lane load fell evenly across the arms, so it does not bias the comparison. The host was loaded (1-min load 4.6-21). The result bounds this head below 9.5% per container. It cannot exclude a rate as low as the one the control showed today. Review 3 (major): nothing would own the ramp quarantine after merge. I searched first: no open issue covered it; the only hit for Every finding
Re-gate on
|
|
Review round 4.2 addressed two minor findings, both on the stall write-up. There is no code change and no new measurement; the head is still 1. Who ran beside the stall containers. Round 4 said another lane's container ran beside every one of the 80. That was wrong. None of the 80 ran alone:
So part of the load that round 4 blamed on other lanes was this lane's own, and that was not disclosed. The own-lane neighbours fell evenly across arms (8 MiB: A 6/10, B 16/30, C 6/10; 128 MiB: A 6/15, B 6/15), so they do not bias the comparison, and the registered verdict stands. While checking, the same misattribution turned up twice more, and both are now fixed:
2. Figures no script regenerated, and the TL;DR. Every stall figure outside the registered
Other changes:
Verdict, unchanged. The stall is EXCLUDED at 8 MiB by the registered rule, with an upper bound of 9.5% per container. The positive control stalled only 1 of 10 times, so the result cannot exclude a rate that low. |
|
Review round 4.3 answered three minor findings. None was on product code. There is no code change: the head is still 1. #708's main baseline came from the loaded re-gate alone (fixed; #708 edited). #708 said main "failed While checking, one slip turned up in 2. A FAIL on FINAL was decided by an instrument that was not registered (disclosed in §6).
No late base re-run was made. On this host, a local 8 MiB run is starved of ring memory by the other root containers (the charge is per UID), and a controlled run would need a ring hog that starves the other lanes. Details: 3. The strict CI gates had one real-CI sample (re-run three times; recorded in §9). Run 36290508684 was re-run three times on the same bytes (attempts 2-4, merge commit
Zero failures in four runs only puts a flake rate below about 53% per run, so this corroborates and does not prove. Census: The merge order is unchanged: this PR first, then #696. Until #696 lands, the |
|
Change of instrument for B1: maintainer decision, 2026-09-27 ~08:45Z. B1 was the pre-registered laptop timing check (S1b run2). It never took a timing sample. Its quiet-host gate never opened: 585 polls from 06:17 to 08:44Z, 1-minute load median 3.3 and never below 1.5, because of interactive use of the host. B1 ended Replacement: a same-run A/B/A′ benchmark on the bare-metal cluster, both arches. Run 36308789830.
B1's witnesses still stand, as supporting evidence (not as data for this run):
The verdict will be posted here when the run completes (ETA ~12:15Z). |
|
Cluster A/B verdict: NO REGRESSION. Run 36308789830 completed successfully at 12:11Z. It was read with the pre-registered
Remaining review minors (the ci.yml ramp-quarantine comment still says main fails the ramps, contradicting the corrected #708; the stall write-up wording) are tracked in a v1.6.0 follow-up under the maintainer's two-round review cap. Merging. |
… hijack accept test #674 (49d2726) added a deferAccept parameter to createListenSocket and updated every caller then on main, passing true to keep the old behaviour (TCP_DEFER_ACCEPT was always set). This PR's TestAcceptOfANumberAHijackReleasedIsOrderedAfterTheHijack was written against the old signature, so after merging main the engine/epoll test package no longer compiled (GOOS=linux go vet: "not enough arguments in call to createListenSocket"). Pass true, as #674 did for its siblings. Checked: GOOS=linux go build ./... and go vet ./... on amd64 and arm64, rc=0 (both failed on the vet before this change).
Fixes #662
Fixes #675
Squash-merge only. Do not merge yet: B1 has no answer (section 10). The standing order for this PR is that it does not merge until a measurement shows no regression on the default configuration. S1 measured it twice and both runs came out INCONCLUSIVE. S1b has not produced a timing result yet (section 10), and its B arm must be this head's product code.
Merge order: this PR first, then #696. #696 (the fix for celeris#691) edits
ci.ymltoo and has to be rebased onto this PR: its new step's comment says theunitjob runs./engine/iouringwithout-v, which stops being true here, and its tests usestartTestEngine, which this PR changes (compatible). Until #696 lands, theengine/iouringstep can go red on celeris#691 at main's own rate; the #657 and #662 steps after it now run anyway (if: ${{ !cancelled() }}), so a red there cannot hide them.Squash message. The repository squashes with the commit messages joined, and this branch carries the abandoned design's commits (
8e3b3ae,810d8c8,10ddfbf,7cd9481,27c07bd). Replace the default with the proposed message at the end of this body.Every file:line below is at the pushed head
7aeb2ec3340bunless another sha is named. Each figure names where it comes from: a script that regenerates it from raw logs, or the report that states it. Section 12 lists which figures are regenerated by one command and which cite a report only. The evidence root,evidence/celeris-662/, is not in this repo.TL;DR
TCP_DEFER_ACCEPTis on (main's default), the kernel keeps a client that has completed its handshake but not yet sent anything out of the accept queue. When a pause closes the listener, that client's first request is answered with a reset, and the engine records nothing: no accept, no close, no error. Every adaptive switch pauses the outgoing engine (epoll: PauseAccept silently drops connections already waiting in the accept queue — their requests get EOF (8/8, deterministic), and a promotion on a GitHub runner loses ~1/3 of 2048 connections #662). So does a standalonePauseAccept(PauseAccept still drops handshake-complete idle connections for standalone engines: TCP_DEFER_ACCEPT stays on by default outside adaptive (the residual of #662, a product call) #675).e.muand never waits for it. Atcp_synack_retries=0guard is added. The round-3switchPossiblegate is deleted.9f4d89b1: RESET 16/16 on both switch directions and 8/8 standalone on both engines. All pass on this branch.read: reset(83 to 392 per ramp-only container).adaptiveTimeoutmerge, survive, and section 8 says why that is expected.engine/iouringstep's good case is real CI at this head.7aeb2ec: run 36290508684, 9 of 9 jobs green on the first attempt and on three re-runs of the same bytes (review round 4.3; merge commit85fb140each time), every tally at its expected count and no--- FAIL:line in any job.1. The loss class and its mechanism
At main
9f4d89b1, both engines setTCP_DEFER_ACCEPTon every listen socket, unconditionally:engine/epoll/loop.go:3118@9f4d89b1, increateListenSocket(:3087@9f4d89b1);engine/iouring/worker.go:5495@9f4d89b1, increateListenSocket(:5468@9f4d89b1).While the option is set, a connection whose handshake has completed but whose client has sent nothing is not in the accept queue. The kernel holds it as a request socket (
TCP_NEW_SYN_RECV), andaccept4answersEAGAIN. #663 added a drain,acceptQueuedOnPause(engine/epoll/loop.go:1054@9f4d89b1,engine/iouring/worker.go:2083@9f4d89b1). It accepts untilEAGAIN, and then the pause closes the listen socket. That close orphans the request socket, and the client's first request is answered with a reset. The engine records no accept, no close and no error.Every adaptive switch pauses the outgoing sub-engine, so every switch is exposed (#662). A standalone epoll or io_uring
PauseAcceptis exposed the same way (#675).The kernel facts the design rests on were read at v6.17 (
evidence/celeris-662/design/DECISION.md§9):net/ipv4/tcp_minisocks.c:860-866@v6.17). Clearing the option therefore stops new deferrals at once.TCP_TIMEOUT_INIT,include/net/tcp.h:155@v6.17). Measured on loopback, promotion took 1002 to 1086 ms (DECISION.md§4.1).syn_ack_recalctakes its non-defer branch. Iftcp_synack_retries=0, it expires a deferred request at the first timer instead of promoting it (net/ipv4/inet_connection_sock.c:872-875@v6.17). Red team S1 lost 256/256 that way; withTCP_SYNCNT=1(S2) it lost 0/256 (DECISION.md§2).The same class seen from the client side is #686. In #687's ramp measurement, 14,554 of 14,554 resets were connections the engine had never accepted. All of them happened at a promote's listener close, and the accept queue was empty at 3,233 of 3,340 closes (
evidence/celeris-657/impl/pr3-r2/measure/MECHANISM.txt). #686 is closed as a duplicate of #662.2. What this PR does: approach C, exactly
Steady state is main's.
createListenSocket(addr, deferAccept)sets the option whendeferAcceptis true (engine/epoll/loop.go:3189,3222-3224;engine/iouring/worker.go:5594,5623-5625). Every caller passes!cfg.DisableDeferAccept, and that field defaults to false. No per-connection code path changes, and the run loops' bodies are compiled to the same instructions as1d90b5d's (section 10).At a pause. Each loop or worker handles the pause on its own thread, at the top of its next iteration (
engine/epoll/loop.go:457-462;engine/iouring/worker.go:1063-1068), instepAcceptPause(engine/epoll/loop.go:1097-1112;engine/iouring/worker.go:2077-2092). It has four transitions:ACTIVE to LINGERING.
deferlinger.Enter(internal/deferlinger/deferlinger_linux.go:41-71)::46-52);:53-62);Lingerafter the setsockopt returned (:70). Red team A3 anchored the deadline at the pause call and lost 245 connections; A4 anchored it at the clear and lost 0 of 4,028 (DECISION.md§4.1).The deadline is on deferlinger's own clock,
Now: nanoseconds since the package's initialisation, read withtime.Since, so on the monotonic clock (internal/deferlinger/deferlinger.go:84-108). It is a distinct type,Deadline, so a deadline cannot be compared with a wall-clock reading without a conversion. A step of the wall clock can therefore neither end a linger early nor keep a paused listener open afterPauseAccepthas returned (review round 4; section 13).Lingeris 1500 ms (internal/deferlinger/deferlinger.go:79). It is internal, and only tests change it (:110-162). A listener that was created without the option, or a linger of 0, closes at once (deferlinger_linux.go:43-45), because nothing on it can be deferred.LINGERING, before the deadline. Nothing changes, and connections are accepted and served through the normal path:
rearmAcceptIfPendingnow gates only onlistenFD >= 0(engine/iouring/worker.go:4600-4605; main's gate atworker.go:4485-4490@9f4d89b1also required!paused).LINGERING to CLOSED, at the deadline.
closeListenerAfterDrainruns. It is main's drain-and-close, moved unchanged into one function (engine/epoll/loop.go:1120-1136;engine/iouring/worker.go:2104-2145). At its end it setslistenFDClosedand wakes a waitingPauseAccept(engine/epoll/loop.go:1134-1135;engine/iouring/worker.go:2143-2144).LINGERING to ACTIVE, on a resume.
deferlinger.Leaveputs the option back on the same fd (deferlinger_linux.go:77-87). The fd is never closed.The
tcp_synack_retries=0guard.BeginPauseAcceptreads the sysctl once, off the loop threads (internal/deferlinger/deferlinger.go:270-282). The pausing listeners getTCP_SYNCNT=1only when that read succeeded and returned 0 (:186-189). An unreadable value never triggers the guard, because the kernel rejectsTCP_SYNCNTvalues below 1: a guard applied on a guess could never be undone. An unreadable sysctl, and a failed clear, guard or restore, are counted every time and logged once per engine (:250-282;deferlinger_linux.go:89-114).How long the loops sleep while lingering.
epoll_waitat the time left to the deadline, including an infinite wait (engine/epoll/loop.go:515-517,:1142-1151). This is defence only: while the listener is open, the wait is already a few ms.adaptiveTimeout = capToDeadline(sweptTimeout(), lingerUntil)while lingering, andsweptTimeout()otherwise (engine/iouring/worker.go:1864-1873,:1892-1894).sweptTimeoutis fix(engine): sweep the connections a switch leaves behind, instead of waiting for each to send again (celeris#657) #687's sweep cap overbaseTimeout(:1877-1888). This is the one design-level merge in the rebase (r6/rebase/00-RESOLUTIONS.md); section 8 has its two mutants.Engine API.
BeginPauseAccept()starts the pause and returns (engine/epoll/engine.go:286-289;engine/iouring/engine.go:565-568). It is exported on the concrete engines only, not added toengine.AcceptController.PauseAccept()isBeginPauseAccept()followed by a wait (engine/epoll/engine.go:305-339;engine/iouring/engine.go:585-619). The wait ends when any of these holds:So
PauseAcceptnow blocks about 1.5 s. It blocks on a per-engine broadcast channel (internal/deferlinger/deferlinger.go:291-345) that a loop's pause close, a loop's exit andResumeAcceptclose, taken before the flags are read so no wakeup is lost, with a 100 ms re-check for the one change nothing notifies (a pause landing on a listener a resume has not re-created yet). It used to poll every millisecond through the linger (section 11).shutdownsetslistenFDClosedand notifies (engine/epoll/loop.go:3167-3168;engine/iouring/worker.go:5528-5529), so a pause does not wait out its bound on a loop that is exiting.Adaptive. Both pause sites in
performSwitchcallbeginPause(adaptive/engine.go:810,:859; helper:909-917). The call is made undere.mu, so any later switch'sResumeAcceptis ordered after it. Nothing waits for the linger. The round-3switchPossiblegate and itsDisableDeferAcceptoverride are deleted (27c07bd). A switch that aborts after building the lazy standby (a driver registered during the build) now also gives the fresh engine the drain a completed switch gives the engine it leaves, toward the engine that stays active (adaptive/engine.go:789-814): what it accepts during its linger moves, instead of staying on a standby no switch may make active.Public surface. One new field,
Config.DisableDeferAccept(config.go:133-143, mapped at:275), which reachesresource.Config.DisableDeferAccept(resource/config.go:70-101). It defaults to false. It is the instant, lossless opt-out: with the option off there is nothing to linger for, and the pause drains and closes at once. There is no new interface method and no new metric. The docs are rewritten for the new behaviour:Server.PauseAccept(server.go:577-598),AcceptController.PauseAccept(engine/engine.go:33-49) andREADME.md:50.3. Why C, and not A or B
The full comparison is
DECISION.md§2. The three candidates were scored the same way:Why not A. A needs no clock, but it fails two of the three correctness scopes, and it fails the release gate by construction:
PauseAcceptstill resets under the default config (PauseAccept still drops handshake-complete idle connections for standalone engines: TCP_DEFER_ACCEPT stays on by default outside adaptive (the residual of #662, a product call) #675: 8/8 on each engine).ForceSwitchon an engine the gate had turned off still resets (census C4:RESET:8).B and C are the same mechanism. C wins on correctness and on API surface.
tcp_synack_retries=0guard and B does not. C's 1.5 s also leaves about 0.41 s of margin over the worst promotion measured on loopback (1.086 s); B's 1.25 s leaves about 0.16 s.DisableDeferAccept, the opt-out that round 3 had already introduced. B would addAcceptPauseLinger(whose negative value re-enables the bug),AsyncAcceptController,ErrPauseWithdrawnand fourEngineMetricsfields.C's own weaknesses were each fixed before implementation (§5 of the record): wait bounds, the guard applied on an unreadable sysctl, the positive-control rule, and existing tests that were silently hollowed out.
4. What ships, file by file
internal/deferlinger/deferlinger.go(new)DefaultLinger1.5 s (:79), the monotonic clock andDeadline(:84-108), test hooks (:110-162),SynackRetriesZero/GuardWanted(:170-189), counters (:193-243),PauseStatewith its once-per-engine warnings and the wait'sChanged/Notify/WaitChanged(:250-359)internal/deferlinger/deferlinger_linux.go(new)Clear/Restore/GuardSyncnt(:14-28),Enter(:41-71),Leave(:77-87),warnOnce(:89-114)engine/epoll/loop.golingerUntil/deferCapable(:121-130); the per-iteration step (:457-472);stepAcceptPause(:1097);closeListenerAfterDrain(:1120);lingerTimeoutMs(:1142);shutdownsets the flag and notifies (:3167-3168);createListenSocket(addr, deferAccept)(:3189)engine/epoll/engine.goBeginPauseAccept(:286),PauseAccept's wait (:305-339),ResumeAcceptnotifies (:344)engine/iouring/worker.go:359-368,:1063-1077,:2077,:2104,:5528-5529,:5594), plus theadaptiveTimeoutcap (:1864-1894) and the re-arm gate (:4600-4605)engine/iouring/engine.goBeginPauseAccept(:565),PauseAccept(:585-619),ResumeAcceptnotifies (:624)adaptive/engine.gobeginPauseat both pause sites (:810,:859,:909-917); the drain on the abort path (:811); comments correctedadaptive/transplant.goswitchWindowHook's windowconfig.go,resource/config.go,server.go,engine/engine.go,README.mdDisableDeferAcceptand the docs (section 2)engine/{epoll,iouring}/pause_accept_linger_linux_test.go(the same line numbers in both), with T1…ServesDeferredIdleConnections:439, T2…ServesArrivalsDuringLinger:643, T3…LingerAnchoredAtClear:682, T4…SynackRetriesZero:454, T5TestResumeDuringLingerRestoresDeferAccept:699, T6TestShutdownDuringLingerIsPrompt:778, and the controls that must lose: N1…NoLingerResetsDeferredIdleConnections:446, N2…NoClearLosesArrivals:666, N3…NoGuardLosesAtSynackRetriesZero:461. Alsopause_accept_idle_linux_test.go(the #662 witness, its no-pause control and the config test),linger_arrivals_placement_linux_test.go(TestLingerArrivalsReachTheIncomingEngine, section 7 T3) andpause_r4_linux_test.go(round 4:TestPauseAcceptWaitIsWokenin both,TestLingerTimeoutNeverBlocksPastTheDeadlinein epoll)engine/iouring/accept_rearm_sqfull_test.goTestAcceptRearmedDuringLinger(:118);TestAcceptRearmRetriedAfterSQFullrewritten for the new re-arm rule (:57)engine/iouring/ring_budget_linux_test.go(new),listen_addr_linux_test.go,driver_test.go(startTestEngineonly), the io_uring linger rigsengine/iouring/ring_budget_linux_test.go:101retryRingENOMEM,:145startRingRetried662; section 7 T2)internal/deferlinger/deferlinger_linux_test.go,deferlinger_r4_linux_test.go(new)adaptive/switch_idle_conns_linux_test.go,adaptive/linger_switch_linux_test.go,adaptive/switch_accept_churn_test.go,adaptive/reverse_transplant_test.go,adaptive/switch_abort_drain_linux_test.go(new)test/deferab/deferab_bench_test.go(new).github/workflows/ci.yml./engine/iouringleaves the root step (:121-125) for its own step (:158-180); the celeris#657 witness step (:227), the celeris#662 linger step (:346-368) and the synack=0 step (:379-402) runif: ${{ !cancelled() }}; the adaptive tally (:551-561); the ramp tests' quarantine points at celeris#708 (:452-464)5. Failing first
Against main
9f4d89b1(r6/gate/logs/g7*.log, regenerated withgrep -HE '^\s*--- (PASS|FAIL|SKIP): ' r6/gate/logs/g7*.log). The runs usedgolang:1.27, memlock 128 MiB,-race -v -count=1, one container at a time. Main's bytes have not changed since (it is still9f4d89b1), so these stand; the branch side is re-run on this head in section 6.A row is a real failing-first witness only if main fails the claim, not just the test's premise. Four rows are; three fail on main only at their premise, because main has no linger to observe. Those three are marked, and the mutant that gives each one a genuine catch on this branch is named (section 8).
9f4d89b1TestSwitchKeepsHandshakedIdleConnectionsoutcomes=map[RESET:16], "the epoll->io_uring switch dropped 16 of 16 clients"…OnRevertoutcomes=map[RESET:16]TestPauseAcceptKeepsHandshakedIdleConnections(this PR's first commit,f0c65e9, written against main's API)AcceptCount = 0, want 8TestPauseAcceptIdleControl, the same rig without the pauseTestPerformSwitchDoesNotBlockOnLingerTestFlapWithinLingerKeepsListenersTestLingerTransplantConservesThree qualifications:
TestAdaptiveListenersKeepDeferAcceptAcrossSwitchespasses on main. It guards against the deleted round-3 gate, which main never had, so it is not a failing-first witness.TestListenSocketDoesNotDeferAcceptfails on main, but it asserts the abandoned first approach (8e3b3ae). It is not counted.BeginPauseAcceptandinternal/deferlinger. Neither exists on9f4d89b1, so these tests do not compile there, and a compile error is not failing-first evidence. What stands in for it is the mutation record (section 8), where each test has a mutant that must make it lose.Round 4's fixes, against the reviewed head
1d90b5d(round4/FF-TALLY.txt,python3 round4/tools/ff-tally.py). Each test is in this head; to run it on1d90b5d, a staging copy that compiles there was dropped in and removed afterwards (round4/ff/, the tree clean before and after:ff/manifest.log). The abort test's copy is byte-identical; the unit tests' copies keep the test bodies and only write1d90b5d's deadline type where it differs.1d90b5dTestEnterDeadlineIsOnTheMonotonicClockTestPauseWarningsAreLoggedOncePerEngineTestLingerTimeoutNeverBlocksPastTheDeadlineTestSwitchAbortedByADriverDrainsTheFreshStandbyTestPauseAcceptWaitIsWoken(both engines)6. G5, the gate that failed on 2026-09-18, and its history
G5 compares full suites, main against branch.
2026-09-18: FAILED at the pre-rebase head
42e5098on main985a386(impl/ship/STOP-DECISION.txt).TestRampH1Asyncat m128 went from PASS to FAIL: main 0/11, branch 11/11 (p = 2.8e-6). With the linger forced to 0, the branch passed 3/3.err=0. So epoll: PauseAccept silently drops connections already waiting in the accept queue — their requests get EOF (8/8, deterministic), and a promotion on a GitHub runner loses ~1/3 of 2048 connections #662's loss was gone, and the placement residue was worse.2026-09-20 10:06:45Z: the prediction was registered publicly, before the rebase (#687, comment 5749130618). Verbatim:
TestRampH1Asyncat m128 passes on fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 rebased onto this PR";unacceptedclass to zero, or near it, on both."2026-09-20 11:08Z-13:21Z: the r6 gate, rebased head
53b52f1against main9f4d89b1(r6/gate/90-GATES.md;cd r6/gate && python3 tools/g5-tally.py). The order was ABBA ABBA. Each log checks that no container ran at its start; none audits starts during the run, and native host load was not recorded.(1) HELD.
TestRampH1Asyncfailed 0/6 on each arm, andTestRampH1Sync0/6 on each. The outgoing epoll held 0 of 2,048 in 36 of 36 high phases per arm.(2) was not tested as registered. The F-vs-S protocol was not re-run. What the gate measured in its place is the ramp tests' own error census, main against branch (
python3 r6/ship/tools/ramp-reset-census.py):read: resetin 13,885,363 requests;That is what (2) predicts, measured by a narrower instrument.
The control still catches the old defect:
42e5098failed 2/2 (995-1,043 left);985a386passed 2/2.2026-09-26: re-gate on
1d90b5d, main alongside: 28 containers, full packages, both memlock shapes, counterbalanced (r6/fix/90-FIX.md). Isolation was from other containers only. The host carried other lanes' nativego testload, which a docker check cannot see (round4/30-CORRECTIONS.md§6).--- FAIL:line.TestRampH1Async(2/2 at m128) andTestRampH1Sync(1/2).2026-09-27: round 4's suites on this head (
round4/SUITE-TALLY.txt,python3 round4/tools/suite-tally.py): the touched packages, full,-race -v -count=1,--cpus 4, m128 and m8. The reference is the re-gate's main and1d90b5dcontainers (main's bytes are unchanged). Other lanes' containers and native processes ran beside these. So did this lane's own stall containers, beside seven of the eight runs atc12d1b8: round 4 ran its suites as a second queue alongside the stall schedule (round4/stall/STALL-CONTEXT.txt§6). Each log records its neighbours and the host load../internal/deferlinger./engine/epoll1d90b5d)./engine/iouringc12d1b8: PASS 180, SKIP 5 (the allow-list), plus the first wake test's race. At7aeb2ecthe local runs were starved of ring memory by other lanes' containers: one run had 1 FAIL at an engine start and 92 SKIP, another rc 0 with 115 SKIP. Real CI at7aeb2ec: green, only the 5 allowed skips./adaptive/...total=0, every high phaseepoll=0of 20481d90b5d's containers, the top-level PASS counts rise by exactly the new tests: +7 in./internal/deferlinger, +2 in./engine/epoll, +1 in./engine/iouring, +2 in./adaptive/...../engine/iouringat 8 MiB, neither a product defect.TestPauseAcceptWaitIsWoken, first version (c12d1b8). It withdrew a pause exactly three re-check periods in and raced the wait's own timer. Rewritten inb32d39f(section 13). Since then it has passed wherever its engine could start: 3 x 3 dev containers, the suites at 128 MiB, the epoll suites, and three times in real CI.TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen/drain_stops_while_held, a celeris#657 test unchanged from main, at7aeb2ec. Its engine start failed withio_uring_setup: cannot allocate memory, 58 ENOMEM lines in that container. On main, that test's helper turns a failed start into "engine did not stop within 5s". The container ran beside another lane's container. Ring memory is charged per UID (r6/gate2/90-GATE2.md, from the kernel source), and every container here runs as root, so one container's rings shrink another's 8 MiB: under the same limit, the same ring hog fit 9 rings once and 6 once (round4/logs/race-hog-*.log). The same package at the same head on a fresh GitHub runner at 8 MiB is green, with only the five allowed skips (section 9). Run again at 8 MiB on7aeb2ec(s2b), the package passed (rc 0), but with 115 skips and 60 ENOMEM lines: other lanes' io_uring containers had most of the 8 MiB. So the local 8 MiB runs of./engine/iouringat this head decide nothing, and real CI stands in for them.round4/r43/10-DEVIATION-AND-CORRECTIONS.md§1).round4/00-PLAN.md§6 says a contemporaneous base re-run decides any FAIL on FINAL that base does not show. None was made (no round-4 log is a9f4d89brun); real CI and the ENOMEM diagnosis decided it instead. The plan's no-push stop could not act either:7aeb2ecwas pushed at 03:07:33Z, before its suites (from 03:09:15Z). The conclusion stands: the test's file is main's byte for byte, base passed the subtest 4 of 4 in the re-gate (2 at 8 MiB), and CI passed it twice at 8 MiB in each of four attempts (section 9). A late base re-run here would be starved of ring memory the same way.7. The three threads the r6 gate left open, and what was done about each
Details:
r6/fix/90-FIX.mdand the round-3 comment on this PR. In short:T1:
TestDriverHTTPZeroOverheadfailed on the branch 2/7 (full package, m8). It is a defect on main, not exposed by this PR.cd r6/fix && python3 tools/t1-classify.py). Pooled, main failed 1/40 and the branch 2/39 (Fisher p = 0.615), which is class D.-count=200per container at m8, main fails more often: 46/600 against 28/600 (Mann-Whitney p = 0.1).UnregisterConnqueues anASYNC_CANCELkeyed by the fd number, and the test closes the fd at once;-EBADF;T2: the two celeris#639 tests skipped at 8 MiB, a silent coverage hole on main that this PR closes.
Ring memory is uncharged 12.4-22.8 ms after a ring closes (
r6/gate2/90-GATE2.md). Engine starts made back to back hit ENOMEM, read it as "io_uring unavailable" and skipped. CI ran the package without-v, so nothing showed.The tests now retry a start that failed only on ring ENOMEM, for up to 10 s (
engine/iouring/ring_budget_linux_test.go:101):engine/iouring/listen_addr_linux_test.go:49,119) andstartTestEngine;startRingRetried662.If io_uring works and the start still fails, the test fails. A ring ENOMEM on the 4-entry probe now counts as usable, and only a positive answer is cached (
engine/iouring/ring_budget_linux_test.go:81-94).Not retried: a start whose constructor's 1-entry capability probe fails.
Newreports that as "io_uring not available on this system", not as ENOMEM. It needs almost the whole budget held, which happened here only beside other lanes' io_uring containers (12 rig skips in each of two starved local 8 MiB runs); the linger step forbids skips, so in CI it would fail, not pass.CI.
./engine/iouringruns in its own-vstep (ci.yml:158-179). The step fails on any skip other than the five that other steps enforce, and requires both celeris#639 tests to pass by name. Real CI:celeris#639 tests want 2, ran 2, passed 2.Not settled, and with no coverage consequence: why the branch's churn test ran slower in gate2.
T3: two interaction mutants survived
TestLingerArrivalsReachTheIncomingEngine. The test gained an ORDER check (engine/epoll/linger_arrivals_placement_linux_test.go:312,429-440; the io_uring twin atengine/iouring/linger_arrivals_placement_linux_test.go:308,427-438). It reads the adopted set first and the outgoing listeners' closed flags second, with no time constant (cd r6/fix && python3 tools/t3-tally.py):Each cell is 2 containers x
-count=5, so 2 independent observations; a mutant's catch is deterministic.8. Mutation record
m1-m11, on
1d90b5d: all caught (r6/fix/G4-TALLY.txt,cd r6/fix && python3 tools/g4-tally.py). Each was restored bycp(RESTORE_OK 24, porcelain 0 after every restore). The code they target is unchanged here, apart fromEnter's return type andPauseAccept's wait.TestPauseAcceptLingerAnchoredAtClear,TestEnterDeadlineAnchoredAtTheClearTestFlapWithinLingerKeepsListenerspausedin the re-arm gateTestAcceptRearmRetriedAfterSQFullperformSwitchTestPerformSwitchDoesNotBlockOnLinger("ForceSwitch took 1.53 s"), the flap testshutdownm12-m18, round 4 (
round4/G4-TALLY.txt,python3 round4/tools/g4-tally.py; applied in a worktree atc12d1b8, whose product code is this head's, and m16 again atb52bc2a, whose wake test is this head's (onlydriver_test.godiffers);round4/mut/manifest.log: every mutant restored bycp, sha256 equal to the HEAD blob, porcelain empty after each). Negative control, unmutated: the linger tests, the deferlinger unit tests and the two new adaptive tests, rc 0, no FAIL, no SKIP.adaptiveTimeoutwithout the linger cap./engine/iouring, m128 and m8, rc 0)adaptiveTimeoutwithbaseTimeoutin place ofsweptTimeoutwhile lingering./engine/iouring, m128 and m8, rc 0)Nowreturnstime.Now().UnixNano())TestEnterDeadlineIsOnTheMonotonicClockTestSwitchAbortedByADriverDrainsTheFreshStandby, m128 and m8Notifydoes nothingTestPauseAcceptWaitIsWoken, both enginesc12d1b8's version of the test; again on the rewritten version, where all five close cycles took 100 ms, one re-check each, and all five resume cycles returned 30-71 ms after the resume, on a re-check)TestLingerTransplantConserves, and 5 epoll linger testslingerTimeoutMsclamp removedTestLingerTimeoutNeverBlocksPastTheDeadlineWhy m12 and m13 survive, and why no test was added for them. Review 2 expected it, from the ORDER check's 1.25 s margin. Neither removes a correctness property; each removes timeliness bounded at 100 ms. While a listener is open,
baseTimeoutnever exceeds its 100 msmaxWait(engine/iouring/worker.go:1897-1926). So without the linger cap (m12) an idle worker closes its listener at most 100 ms after its deadline, still accepting and serving until then. Without the sweep cap while lingering (m13), the sweep's next pass over a linger arrival is at most 100 ms away instead of on its own cadence, and every linger arrival is still placed before the first close. A test tight enough to see a delay under 100 ms would be a timing test on a CI runner, and neither mutant exposes a defect, so the plan's rule (round4/00-PLAN.md§5) says report, not add.9. CI: the interlocks, proven both ways, and the real run
Interlocks, round 4 (
round4/CI-INTERLOCK.txt,cd round4 && python3 ci/classify.py). Round 4 changed the run blocks of the celeris#662 linger step (two names added) and the adaptive step (two names added), and putif: ${{ !cancelled() }}on the #657 witness, linger and synack steps. Each step was extracted verbatim fromci.ymlwith PyYAML (ci/00-extract-steps*.log, whose manifest records each step'sif:as parsed) and run the way GitHub runs it (bash --noprofile --norc -eo pipefail, the step's own env, a container at the runner's 8 MiB).unit)want 28, ran 28, passed 28, SKIP lines 0); every bad input rc 1, at the tally, and m4 atgo testwant 4, ran 4, passed 4); every bad input rc 1, at the tally, m5 atgo testwant 9, ran 9, passed 9, SKIP lines 0); every bad input rc 1b32d39f: good rc 0 (SKIP lines 5, not enforced by another step 0; celeris#639 tests ... passed 2). The other-test-skipped case went red (rc 1) but atgo test, before the tally: thestartTestEnginedata race (fixed inb52bc2a) plus ring ENOMEM, beside two of this lane's own io_uring stall containers (128 MiB) and one other lane's container (round4/stall/STALL-CONTEXT.txt§6). At7aeb2ecthe local runs were starved of ring memory by other lanes' containers (64 ENOMEM lines, 118 skips), and the step went red, as it must when tests skip. The clean good case is real CI at7aeb2ec: rc 0,SKIP lines 5, not enforced by another step 0; celeris#639 tests want 2, ran 2, passed 2. The step's bad inputs (absent, renamed, skipped, subtest skipped, another test skipped, m639, m639r) were proven in r6 on this byte-identical run block./engine/iouringis not among its packagesif: ${{ !cancelled() }}itself is GitHub's step condition, which a container cannot evaluate. What was checked: the parsed YAML carries it on the three steps (ci/00-extract-steps*.log), actionlint accepts it, and the real run below ran every step.The r6 interlock proof (28 of 28 cases,
r6/fix/CI-INTERLOCK.txt) covers theengine/iouringstep's other cases; its run block is byte-identical at this head.Real CI on
7aeb2ec. Run 36290508684 (pull_request): 9 of 9 jobs succeeded on the first attempt. The job logs were fetched withgh api …/actions/jobs/<id>/logs --allow-escape-sequencesintoround4/ci/github-run-36290508684/and tallied bybash round4/tools/ci-census.sh(round4/ci/CI-CENSUS-36290508684.txt). There is no--- FAIL:line in any job. The only--- SKIP:lines are the five on theengine/iouringstep's allow-list. The runner was Ubuntu 24.04, linux/amd64, kernel6.17.0-1022-azure. The tally lines, verbatim:The round-4 tests there:
TestPauseAcceptWaitIsWokenpassed three times: in theengine/iouringstep at 8 MiB, and in the linger step for both engines. At 8 MiB it logged "workers=1 default linger: pause took 1.5s, woken=1 rechecks=14".TestSwitchAbortedByADriverDrainsTheFreshStandby: 23 of 32 accepted during the linger, 0 left.TestLingerTransplantConserves: keep-alive err=0, churn err=0.maxAcceptranged from 40 µs to 810 µs, against its 100 ms bound.Re-runs of the same bytes (review round 4.3). Until then, one green run was the only real-CI sample of the new strict gates (the
engine/iouringskip allow-list and the adaptive tally of 9), and the local 8 MiB runs cannot stand in for it (section 6). So run 36290508684 was re-run three times (attempts 2-4, 04:53Z-05:16Z), each on the same merge commit85fb140. All four attempts were 9 of 9 green, and each time:DATA RACE;engine/iouringstep took 155.7-156.7 s of its 300 s, anddrain_stops_while_heldpassed twice at 8 MiB;TestLingerTransplantConserveshad err=0 on both counts (839-1,584 transplanted, all adopted), and the default-linger storm passed (its dial errors, logged but not asserted: 1, 0, 1, 0).Zero failures in four runs bounds a per-run flake rate only below about 53% (one-sided 95%). A flake would show as a red run, because every tally needs its exact count (
round4/r43/CI-RERUN-CENSUS.txt,bash round4/r43/tools/ci-census-rerun.sh).Build and lint (
round4/logs/g1g2/; the same containers, pins and controls as r6/fix):go buildandgo vetpassed on all five modules for linux/amd64 and linux/arm64 (10 of 10 each), andgo test -con the five packages this PR touches, both arches.gofmt -lnames onlytest/benchcmp_ws/bench_test.go(pre-existing, section 11).unix.SetsockoptInt(errcheck, LINT_RC 1), an unpinnedactions/stale@v9(unpinned-uses, rc 14), and an unquoted$(mktemp)variable in theengine/iouringstep (SC2086 on line 163 of the planted copy'sci.yml, rc 1). The actionlint control first ran on a copy without.gitand exited 3 ("no project was found"), which is not a detection. With an emptygit initit reported the plant (logs/g1g2/g2-actionlint-poscontrol-git.log).10. B1: no measured regression on the default configuration — NOT ANSWERED
This PR does not merge until B1 is answered by measurement. That is the maintainer's standing order. "Instrument-limited" is not an accepted answer.
S1 (2026-09-18,
impl/s1/) measured42e5098against985a386, 40 blocks per shape. Primary and replication each gave 12 EXCLUDED, 0 REGRESSION, 4 INCONCLUSIVE (ns/opon ChurnSilent and KeepAlive). No estimate reached +2%.It was inconclusive because of the instrument, not the effect (
s1b/ANALYSIS.md):S1b (2026-09-27) stopped before its first timed round (
s1b/run/). It was pre-registered against1d90b5d. Its static witness (E2) found one more changed function than registered (the base's adaptive: forced switches leave keep-alive connections behind — async conns miss the 1.2 s flap window in every run, and a one-worker io_uring strands sync conns until their requests fail #657 sweep) and a steady-path delta of +5 instructions where +4 was registered. Both are explained statically, but the registered rule stopped the run. Its socket witness held 8/8.What round 4 changes for B1. Round 4 changes product code, so B1's B arm must now be
7aeb2ec. It does not change the steady path. A per-function comparison of the compiled test binaries of1d90b5dand this head (round4/static/STATIC-WITNESS.txt,python3 round4/static/diff.py, code addresses and global-page offsets masked) finds(*Loop).runand(*Worker).runidentical instruction for instruction:The product functions that changed:
performSwitch(the abort drain);PauseStateembedded in each engine grew.So S1b's steady-path expectation carries over from
1d90b5d, and its list of changed functions has to add these.Any REGRESSION cell blocks the merge. If S1b is still inconclusive, the result goes back to the maintainer.
11. Known residuals, and the keep-alive stall
The keep-alive stall at a one-worker io_uring revert (review round 4's major;
round4/stall/STALL-TALLY.txt,python3 round4/stall/analyze.py). At42e5098a keep-alive connection sometimes got no answer to a request it sent just after a revert (io_uring -> epoll) with one io_uring worker: 8 of 30 rounds in the D2 mechanism probe, and main stalled once too (impl/price/99-PRICE-SUMMARY.txt, S2b/D2). It was neither attributed nor excluded, and the linger raises the traffic on exactly that path (every connection the outgoing engine accepts during its linger is transplanted: 1,459 to 1,622 per revert against 28 to 32 on main). Round 4 measured it on this head's product code against main, with D2's own rig and schedule, pre-registered before the first container (round4/00-PLAN.md§3, hashed): 5k conn/s churn plus 32 keep-alives at one request per 10 ms, four forced switches per container, one container per observation, counterbalanced, memlock 8 MiB (one io_uring worker, the CI shape) and 128 MiB (four).Arms: A = main
9f4d89b1; B = this head's product code (built atc12d1b8; every later round-4 commit changes tests only); C = the D2 binary of42e5098itself, sha-pinned, as the positive control. 80 containers, 2026-09-27 01:44Z-03:05Z, all 80 valid; in A and B every switch to io_uring recorded 1 worker at 8 MiB and 4 at 128 MiB.42e5098(control)The counts, the bounds and p = 0.25 are the registered analysis in
STALL-TALLY.txt, and the churn outcomes in the last paragraph are that file's descriptive part. Every other stall figure below, including the C row's detail, is regenerated bypython3 round4/stall/context.pyintoround4/stall/STALL-CONTEXT.txt, which was added in review round 4.2. None of those figures was registered.By the pre-registered rule, the stall is EXCLUDED at 8 MiB on this head. The control that let the rule decide was weak. The rule asks only that C stall at least once, and it did. B stalled in 0 of 30 containers, so the one-sided 95% upper bound on B's rate is 9.5% per container, 4.9% per revert (60 reverts). At 128 MiB, 0 of 15 gives 18.1% per container, 9.5% per revert. The only comparison the plan registered is B against today's control: 0 of 30 against 1 of 10, one-sided Fisher p = 0.25, which is not significant. Main did not stall either (0 of 25).
Cross-session comparisons (not registered;
STALL-CONTEXT.txt§5, with D2 recomputed from its raw rounds). C is byte-identical to the binary that stalled 8 of 30 times in D2. Against that rate, B's 0 of 30 is lower at one-sided Fisher p = 0.0023, and against today's C and D2 pooled (9 of 40), p = 0.004. Today's 1 of 10 and D2's 8 of 30 do not differ significantly (two-sided p = 0.40), so the data do not show that the control stalled less today. The cross-session figures also assume the two sessions are comparable, and the conditions below limit that.Conditions (
STALL-CONTEXT.txt§1-3, corrected in round 4.2).-racesuites, the G4 mutants and the CI-interlock cases. While one of those ran beside a stall container, the lane held both slots itself.Round 4 said that another lane's container ran beside all 80. It blamed the weak control on other lanes' load and did not disclose that part of that load was this lane's own. The own-lane neighbours fell evenly across the arms (8 MiB: A 6 of 10, B 16 of 30, C 6 of 10; 128 MiB: A 6 of 15, B 6 of 15), so they do not bias B against A or C. The one control stall had only another lane's container beside it.
The host's 1-min load at the containers' starts and ends was 4.6-21 (median 7.6). D2's sampler (
impl/price/swdiag/sampler.log, every 5 s, 247 samples with a container running) saw only D2's own containers, at load 1.2-4.0 (median 2.3). Round 4 gave that range as 1.2-3.4, but the sampler's maximum is 3.97.What the result bounds is this head's rate under today's conditions: below 9.5% per container at 8 MiB. It cannot exclude a rate as low as the one the control showed today. Descriptively, the churn connections that main lost at a listener close (
read: reset, the #662 class) were 27 at 8 MiB and 54 at 128 MiB. This head lost 1 in all 45 containers, plus 4 resets at dial.Inherent to closing a listen socket, and unchanged (at
42e5098,impl/price/99-PRICE-SUMMARY.txt, E3):EAGAINand the close.On loopback with arrivals every 2 ms, this branch lost 0 of 409 at the close and 0 of 72,077 during the linger. Main lost 311/311 and 125/125.
Lossy paths (E4, netem, at
42e5098). With no packet loss, 0 of 2,048 silent clients were lost at every RTT up to 400 ms. At loss p the branch loses about 1 − (1 − p)² of them, because both the retransmitted SYN-ACK and its answer must arrive: 2.0-2.4% at 1% loss. The opt-out loses about p. Main loses 100%. Not covered: non-Linux clients, middleboxes, a BPFTCP_TIMEOUT_INIToverride.The price of the pause.
PauseAcceptnow blocks about 1.5 s (epoll p95 1503.4 to 1505.0 ms; io_uring idle p95 1526 to 1536 ms, at42e5098, S2a), and keeps admitting connections during that time.PauseAcceptused to spend CPU in its own wait: it slept 1 ms at a time through the linger, and on an idle io_uring engine every wakeup woke a parked thread. D1 measured 65-73 ms per pause at42e5098. Round 4 fixed that and measured it before and after, pre-registered (round4/00-PLAN.md§4;round4/cpu/CPU-TALLY.txt,python3 round4/cpu/analyze.py), with D1's rig: 3 containers per head per shape, 4 cycles each. Per container, the median CPU of the wait (D1's c3 - c2) was, on io_uring, 43.1 to 66.5 ms on1d90b5dand 0.8 to 4.0 ms on this head; on epoll the medians per shape went from 16.4 and 16.6 ms to 2.0 and 5.1 ms. Two of the three registered predictions held (P1, P3). P2, the pause's wall time within ±10 ms, missed at 8 MiB: +13.1 ms (median, this head minus1d90b5d); at 128 MiB it was -4.5 ms. That wall time is set by when an idle io_uring worker notices the pause (up to 100 ms) plus the 1.5 s linger, not by the wait. On1d90b5dalone the per-container medians at 8 MiB range from 1504 to 1523 ms.SO_REUSEPORTshare of new connections for 1.5 s (0.50 at m128). Everything it accepts during that time is transplanted: 5,948 to 6,126 connections per promotion, against 28 to 32 on main (S2b, at42e5098).Other open items.
-skip(ci.yml:520), owned now by CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708, which records what lifting them takes (at least 6 GitHub-hosted runs of identical bytes, then a PASS tally). They pass locally on this head at m128 (round 4's adaptive suite:TestRampH1SyncandTestRampH1AsyncPASS, error censustotal=0), as in section 6. CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708's baseline was corrected in round 4.3: it had pooled main's record from the loaded re-gate alone, including P7's contaminated main arm. It now shows the r6 gate separately, where main passed both ramps in 6 of 6 containers (round4/r43/RAMP-BASELINE.txt).engine/iouringstep red at main's own rate until fix(iouring): run every driver op through the engine's own duplicate of the socket, and count every cancel until its CQE, so closing after UnregisterConn is safe (celeris#691, celeris#707) #696 lands. The steps after it now run anyway.Workers: 1; they pass with 2), and the root package's celeris#592 io_uring subtests skip at 8 MiB.TestWriteBufBackpressureClosesSlowConsumeris opt-in by design (noted in CI: two test groups have never run — the epoll sendfile e2e tests (Workers: 1 fails validation, the helper skips) and the #592 settled-route io_uring subtests (skip at 8 MiB) #709). Separately,test/benchcmp_ws/bench_test.gois not gofmt-clean; it has been unchanged since v1.4.2, and no CI lint job covers that module.12. Reproduce
Base dir:
evidence/celeris-662/.Regenerated from raw logs by one command each.
1d90b5d:bash r6/ship/tools/regen-all.shwritesr6/ship/REGEN.txt. It covers the G5 tallies of the r6 gate and the re-gate, G4 (m1-m11), T1, T3, CI-INTERLOCK (r6), T2c, the S1 per-cell table, the ramp census and the CI census.t2b/tools/tally-void.pyregenerates its tally with the void containers marked;t2b/T2B-TALLY.pre-void.txtkeeps the tally as first published.r6/round4/MANIFEST.txt):python3 stall/analyze.py(the stall, the registered analysis),python3 stall/context.py(the stall section's other figures, none registered: neighbours, host load, the control stall's ledger, the cross-session comparisons with D2; added in round 4.2),python3 cpu/analyze.py(PauseAccept CPU),python3 tools/suite-tally.py(the suites),python3 tools/g4-tally.py(m12-m18),python3 tools/ff-tally.py(failing-first against1d90b5d),python3 ci/classify.py(the interlocks),python3 static/diff.py(the compiled run loops).r6/round4/r43/, run fromr6/round4):python3 r43/tools/ramp-baseline.py(CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708's baseline, per container, with the r6 gate and the re-gate kept apart) andbash r43/tools/ci-census-rerun.sh(every attempt of run 36290508684).Cited from a report, not regenerated by one command here. Each report names the scripts that produced it, beside it:
design/DECISION.md§§2, 4.1, 9;design/DECISION.md§§2, 4.1;r6/gate2/90-GATE2.md, tallied byr6/gate2/tools/t2-tally.py;42e5098:impl/price/99-PRICE-SUMMARY.txt, regenerated by the scripts inimpl/price/tools/.Every container log records its head sha,
git status --porcelainanddocker ps; round 4's also record the shared slot, the containers that started during it, and the host load.13. Review round 4: what changed on top of
1d90b5dThree reviews of
1d90b5d(one APPROVE, two REQUEST_CHANGES) raised two majors and a set of minors and nits. The round-4 comment on this PR answers each one. The commits, fast-forward on1d90b5d:30ba519,85545b2): the linger deadline on the monotonic clock; a standalonePauseAcceptwoken by the pause close, a loop's exit or a resume instead of a 1 ms poll; each pause warning logged once per engine; the epoll linger cap applied to an infinite wait; and an aborted switch drains its fresh standby. Each has a test that fails on1d90b5d(section 5) and a mutant (section 8). RULE 10: the one new lock is a leaf, and the abort drain runs where the completed switch's already does (round4/20-DEADLOCK.md). RULE 74: the run loops compile to the same instructions (section 10).551d62f,71cc7a0): the linger rigs' goroutines stop on an early failure; the io_uring linger rigs retry a ring ENOMEM start;ioUringUsableHereno longer caches a transient ENOMEM; T2 logs its accept margin;TestLingerTransplantConservesfails on a lost request; the switch storm runs at the default linger too.c12d1b8): the ramp quarantine points at celeris#708; the adaptive: forced switches leave keep-alive connections behind — async conns miss the 1.2 s flap window in every run, and a one-worker io_uring strands sync conns until their requests fail #657 witness, epoll: PauseAccept silently drops connections already waiting in the accept queue — their requests get EOF (8/8, deterministic), and a promotion on a GitHub runner loses ~1/3 of 2048 connections #662 linger and synack steps run after a red step; the tallies name the new tests.b32d39f,b52bc2a,7aeb2ec, tests only):TestPauseAcceptWaitIsWokenwithdrew its pause exactly three re-check periods in, raced the wait's timer, and failed once in the./engine/iouringsuite at 8 MiB. It now asserts the wakeup over five pauses per path, where only its total absence fails (and m16 is still caught every time);startTestEngine's ENOMEM retry, added in9a01f71, shared onedonechannel across attempts. A CI-shape run in which starts were retried under-race(theiouringpkg-otherskipinterlock case, with two of this lane's own io_uring stall containers and one other lane's container beside it) reported seven data races, each failing the test whose start was retried. Each attempt now has its own. Shown both ways with a ring hog that makes the first attempts fail (round4/logs/race-hog-*.log): before, one DATA RACE atengine/iouring/driver_test.go:66,70@c12d1b8and a FAIL; after, 199 retries over 2.5 s, no race, PASS.startTestEnginenow logs when it retried.Review round 4.3 (three minor findings, no code change): #708's baseline was split into the r6 gate and the loaded re-gate (section 11); the section-6 deviations were disclosed (
round4/r43/10-DEVIATION-AND-CORRECTIONS.md); and CI was re-run three times on the same bytes, all green (section 9).Proposed squash commit message
Title:
Message:
Summary by CodeRabbit
DisableDeferAcceptoption to close listening sockets immediately when pausing acceptance.