fix(iouring): hand a connection to epoll only when no read can still resolve its fd (celeris#657) - #681
Conversation
…etime (celeris#657) PR 2 of 3 for #657 (face 2). These tests fail on main (5b2e83b) and pin what the fix must do; the fix follows in separate commits. Unit tests drive one connection through the real request, send and hand-off code with the kernel taken out of the loop: every SQE the engine places is read back and removed unsubmitted, and every completion is written by the test, so each ordering a cancel can resolve in is a deterministic input. - TestTransplantNeverHandsOffArmedRecv: no hand-off while the recv is armed; a reported cancel of its own tag (0x09) instead, hand-off at the recv's -ECANCELED, a request that beats the cancel served and HELD, and the same against the real kernel. - TestTransplantReapMissIsRetried: a missed reap is retried, never followed by a hand-off; a stale miss cannot clear a newer reap. - TestHoldReleasedWhenDrainStops: a held conn gets its recv armed at the SEND completion when it is not handed off (processCQE, the worker loop, a refused hand-off). - TestHoldRescuedByCheckTimeouts: the timeout sweep's belt, counted. - TestOneOwnerPerHandoff: a conn claimed by its dispatch goroutine is not also moved by tryTransplant; finishAsyncTransplant leaves a slot it no longer owns alone. - TestNoDrainSQESequenceIsUnchanged: with no drain set the per-request SQEs (opcode, flags incl. IO_LINK, tag) of the sync, WRITEV and direct-body tails. Passes on main; it is the witness that the fix changes nothing on the steady-state path. Engine-level (real worker loop, keep-alive clients across a StartTransplant): TestHandoffHasNothingInFlight, TestStaleRecvDataCounted (client losses == stale data CQEs, and both 0), TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen and TestWorkerParksWithNothingPending. Refs #657
…sites (celeris#657) tryTransplant and finishAsyncTransplant each carried their own copy of the commit sequence (dup, detach, witnesses, cancel, release, close, adopt). Both now call handOff once their gates have passed, so the fd-lifetime gate that follows is written once and the celeris#657 witnesses stay where a gate regression would move them. No behaviour change: the order of every step and the release kind per site are as before. Refs #657
…celeris#657) EngineMetrics gains five io_uring-only, cumulative fields, exported the way the PR-1 witnesses are (the engine's handoffLoss set, Metrics(), and the adaptive sum over both sub-engines): - TransplantHeld: responses flushed with the next recv held for a hand-off (HOLD). - TransplantReaps / TransplantReapMisses: reported cancels of an armed recv submitted so the conn can be handed off (REAP), and those that matched nothing. - TransplantHoldRescued: held conns the timeout sweep had to rescue. Must stay 0. - TransplantDoubleClaim: hand-offs refused because another path owned the conn's hand-off (a double hand-off prevented). Nothing increments them yet; the rule that does follows. Refs #657
…ve its fd (celeris#657) P1 (R0 gate) and P2 (REAP) of the fd-lifetime rule. Both hand-off sites, tryTransplant and finishAsyncTransplant, now refuse a connection with recvArmed, kernelInflight != 0 or zcNotifPending. When the only op in flight is the recv, they reap it: a REPORTED cancel matched on the recv's own user_data (generation included) under a new tag, udTransplantReap = 0x09<<56. The hand-off runs again at the recv's -ECANCELED, which handleRecv now routes first, before the generic negative-result branch that closes the conn. finishAsyncTransplant reaps instead of cancelling, dup'ing and closing with the recv armed. - If the request arrives first, the recv completes with it and it is served here as usual. - A cancel that reports a miss is retried on the next loop iteration while the recv is still armed (drainDetachQueue runs the retries). A miss is never followed by a hand-off. - The per-connection state is a COUNT of outstanding reaps plus a flag saying they all target recvs that have since completed, so a newer recv can get its own reap and a stale miss cannot clear it (the celeris#484/#596 lesson). - A pause's cancel keeps precedence in the -ECANCELED routing; detached WS/SSE conns are never reaped. The PR-1 witness tests drove the sites with a recv armed, which the gate now refuses. They check the refusal and then drive handOff, the commit point the gate guards, so the witnesses stay proven. Refs #657
…s set (celeris#657) P3 (HOLD). While a drain is set and there is no provided-buffer ring, the response of a connection the hand-off would accept is flushed unlinked and its next recv is not armed at all. That SEND's completion then finds nothing in flight and hands the connection off; the client's next request waits in the socket buffer for epoll. Every recv-arming tail after a response now goes through one helper, respondAndArm: the sync tail (flushSendLink, or flushSend on a buffer ring) and the direct-body tail (flushSend, then a standalone recv). With no drain set it places exactly the SQEs the two tails placed before (TestNoDrainSQESequenceIsUnchanged); the added cost per request is one atomic pointer load. Held connections the hand-off does not take are released by the next commit. Refs #657
…e (celeris#657) P4, the release guarantee for HOLD. A held connection has no recv armed, so if the hand-off at its SEND completion does not happen (the drain stopped, a gate refused, the dup failed) nothing would ever read its next request. - releaseHold runs after the transplant attempt at both inlined dispatch sites, unconditionally: once nothing is left to send, the hold is cleared and the recv armed. - processCQE's udRecv and udSend cases gain the transplant attempt they lacked, and the release. The listener-close harvest handles completions there. - checkTimeouts rescues a held connection found with its send done and arms its recv, counted in TransplantHoldRescued, which must stay 0. Refs #657
P5 (A6). A promoted async conn whose dispatch goroutine had parked and claimed its own hand-off (transplantPending, asyncRun cleared) could be moved a second time: a SEND completion landing before the worker drained the claim reached tryTransplant, which only skipped a RUNNING goroutine and moved the conn as if it were sync; drainDetachQueue's finishAsyncTransplant then dup'd cs.fd, a number the first move had closed and the process had reused. Measured on main: one identity moved twice, 20 us apart, in 1 of 167 runs, joined to the only negative gauge. - tryTransplant reads the claim with asyncRun under asyncInMu (the lock the goroutine sets both under) and refuses a claimed conn. - finishAsyncTransplant returns unless w.conns[cs.fd] == cs and the conn is not closing. Both refusals count in TransplantDoubleClaim. Refs #657
P6 (A5). The iteration that closes or hands off a worker's last connection queues that connection's close-path cancels in the same pass that finds the worker idle, and the DRAINING->SUSPENDED park is indefinite, so the worker parked with them unsubmitted (measured 1-24 pending SQEs at parks). It now submits before taking wakeMu and setting suspended. Refs #657
…ting target (celeris#657) The warm-up response's SEND completion can reach the worker after the client has read the response, and so after StartTransplant. The idle conn is then reaped and handed to the target before /big is sent; the test's accepting target closed it, and the client read a reset (8 of 15 runs on 17ab243, reported with adopted=1 held=0). The target now refuses, so such a conn is reclaimed onto io_uring and /big is served and held there as the test intends; it also asserts the response really was held, and reports the hand-off counters when the read fails. Refs #657
…en, on one worker (celeris#657) The unit job runs ./engine/iouring without -v, so a skip there is invisible. The ten fd-lifetime tests run again by name with -v and CELERIS_REQUIRE_IOURING_WORKERS=1, with the witness step's exact tally (RUN and PASS counts, no SKIP line). It is also the one-worker leg: the runner's 8 MiB memlock funds a single io_uring worker, the shape the loss was measured in, and the step fails unless every engine the tests start logs workers=1. Refs #657
…fd-lifetime counters (celeris#657) TestHandoffHasNothingInFlight gains an async subtest: every route async, so each conn is promoted to a dispatch goroutine, claims its own hand-off at its park and leaves through finishAsyncTransplant, which must reap the recv the feed path armed. The load helper logs the new counters (held, reaps, misses, rescued, double claims) next to the witnesses. Refs #657
…n the revert tests on one worker (celeris#657) The unit job's celeris#657 witness step now carries all twenty-one tests: the eleven hand-off loss witnesses from PR-1 and the ten fd-lifetime tests this PR adds. It runs them by name with -v, CELERIS_REQUIRE_IOURING_WORKERS=1 and an exact tally (twenty-one RUN, twenty-one PASS, no SKIP line). Each engine-level test logs its io_uring worker count, and the step requires every such line to read workers=1 at the runner's own 8 MiB memlock, the one-ring shape in which the hand-off losses were measured. The adaptive job gains a second step for TestReverseTransplant and TestBidirectionalFlap at the runner's memlock. The job's existing step raises memlock, so until now those two tests never ran with io_uring on a single worker. The new step forbids skipping (neither test reads a REQUIRE variable, so the tally fails on any SKIP line), requires exactly two RUN and two PASS, and requires every "io_uring engine listening" line to read workers=1. It runs even when the step before it failed.
…plit the double-claim counter (celeris#657) Round 2 of #681 (PR 2 of 3 for #657). Review findings C1-C4, C6 and the m15 survivor. C1. The reap cancels an armed recv with IORING_ASYNC_CANCEL_ALL, and cancel flags exist only from Linux 5.19. Through 5.18 the kernel fails the cancel with -EINVAL and leaves the recv armed (measured on 5.15.0-191), and handleTransplantReap read -EINVAL as a miss and requeued it: one failing cancel per loop iteration, per connection, for as long as the drain lasted. - New probeAsyncCancelFlags (cached, like the other runtime probes) submits exactly the reap's SQE form against a user_data nothing carries: res >= 0 or -ENOENT means accepted, -EINVAL rejected. New stores the answer, logs it (async_cancel_flags), and createWorkers copies it to every worker. - startReap places no reap where the flags are rejected and queues no retry. The connection keeps its recv until the recv completes on its own; a sync connection is then HELD at its next response and leaves at that SEND's completion with nothing in flight. Counted in the new TransplantReapUnsupported (a rate, 0 from 5.19). - handleTransplantReap retries only a miss (0, -ENOENT). Any other non-hit result is counted in the new TransplantReapFailed (must stay 0) and is neither retried nor followed by a hand-off. The other cancel sites have the same pre-5.19 problem on main; they are not changed here. C2. A hand-off that fails at its dup (EMFILE) left the conn in place, and the recv re-armed for it was reaped again at once: a RECV, a cancel and two completions per iteration while the failure lasted. handOff now sets reapSuppressed on a dup or SetNonblock failure; startReap places nothing while it is set; handleRecv clears it when data arrives. The dup goes through a Worker.dupFD test seam (nil = unix.Dup). C3. TransplantDoubleClaim counted two different things. tryTransplant finding a conn whose dispatch goroutine had claimed its own hand-off is ordering (a completion landed between the park and the drain of the claim); it is now TransplantClaimDeferred, a rate. finishAsyncTransplant counts TransplantDoubleClaim only when the connState no longer owns its slot, which in async mode means another hand-off of the same conn; a closing conn and the range check are plain refusals. So a release gate of TransplantDoubleClaim == 0 now means what it says. C4. reapOutcome consumes the reaped recv's last completion, so it now does what handleRecv does at any recv's last completion: it returns a flagged provided buffer to the ring and clears recvLinked. C6. The A5 comment said SQPOLL "submits by itself"; an idle SQ thread needs the NEED_WAKEUP kick. Comment corrected; no tier enables SQPOLL. m15 (removing releaseHold after the inlined udRecv dispatch) survived round 1's mutation run. It is an equivalent mutant, so the call is removed, together with the same call in processCQE's udRecv case, which has the same status (round 1's m12 was killed only through the send site). Proof, from code at 86d46ac: - transplantHold is set in one place, respondAndArm, only with no provided-buffer ring (single-shot: the recv that carried the request has just completed, nothing is armed), and only when holdEligible's last clause holds (cs.sending or bytes pending). flushSend never completes a send synchronously. - Every recv-arming path clears the hold first (releaseHoldSlow, rescueHold) or refuses a held conn (reapOutcome), and detached, paused, H2 and promoted async conns are never held. - So a held conn has no recv armed, and the only recv completion that reaches a recv-site release with the hold set is the one whose own handleRecv just set it, with its SEND pending: releaseHoldSlow returns at its egress check without changing anything. - A closing conn only has the flag cleared, which nothing observes (checkTimeouts skips closing conns; releaseConnState resets it). Measured (86d46ac plus a counting overlay, -race, 8 MiB, one worker): the recv-inline site was entered 60,559 times with the hold set and changed nothing, processCQE's recv site 6 times, nothing; in the same runs the send-inline site released 86 of 90 entries (the positive control). The send-site releases and the checkTimeouts rescue stay. Tests (engine/iouring): TestTransplantReapFailureIsNotRetried, TestNoReapWithoutAsyncCancelFlags, TestReapSuppressedAfterFailedHandOff, TestReapedRecvLeavesNoLinkOrBuffer, TestAsyncCancelProbeClassifies, TestAsyncCancelProbeOnThisKernel, TestWorkersCarryTheAsyncCancelProbe and TestReapOnTheRunningKernel; TestOneOwnerPerHandoff gains the split and a closing-conn case; kernel_cancels_the_armed_recv also checks a kernel that rejects the flags. The new EngineMetrics fields are summed by the adaptive engine and pinned by both metrics tests.
…off loss (celeris#657) TestFlapConnsPerRing fixes both engines at two workers, so 256 keep-alive connections are 128 per io_uring ring whatever memlock funds, and runs three promote/revert cycles under continuous load, requiring zero client errors. Step 0 validated it as the CI gate for face 2 (main FAIL 6/6), and the #681 campaign measured it: base 5b2e83b FAIL 8/8 (1,713 lost requests, each one a stale recv the witness counted), fix PASS 8/8. Until now it existed only as an overlay. The file is the measured step-0 source with four changes: its headers no longer say "overlay only", the conditional six-cycle variant (never triggered) is dropped, the two files are one, and a premise guard skips it when RLIMIT_MEMLOCK funds fewer than two io_uring workers, failing instead under CELERIS_REQUIRE_UPSWITCH=1 as the adaptive CI job sets it.
…r leg, and run the new fd-lifetime tests (celeris#657) - adaptive race step (memlock unlimited): now shell: bash with pipefail and a tee, plus a tally for TestFlapConnsPerRing: exactly one top-level RUN and PASS, no SKIP line for it, and exactly one RESULT line showing both engines at two workers and err=0. At raised memlock the two revert tests did not lose on the base; T4 did in 8 of 8 runs. - one-worker revert leg: the verdicts of TestReverseTransplant and TestBidirectionalFlap tolerate up to 64 lost requests (a base flap run PASSED with 56), so the step also requires each test's own summary line (async=false) to read err=0. - unit witness step: the eight new fd-lifetime tests join the PR-2 list, 29 in all.
…load line (celeris#657) The celeris657 load line of the engine-level fd-lifetime tests now also prints TransplantClaimDeferred, TransplantReapFailed and TransplantReapUnsupported (read by name, -1 on a tree without them), so a run on a kernel without cancel flags shows which path its connections took.
… and pin that data lifts the reap suppression (celeris#657) - probeAsyncCancelFlags is probeAsyncCancel(IORING_ASYNC_CANCEL_ALL). TestAsyncCancelProbeOnThisKernel also submits a cancel_flags bit no kernel defines, which every kernel rejects with -EINVAL, so a probe that stopped reading the kernel's answer fails on a 5.19+ kernel too. - TestReapSuppressedAfterFailedHandOff now ends with the conn receiving data again, being served with the drain stopped, and having its linked RECV reaped once the drain is back: the suppression lasts until data, not forever.
|
Round 2 is pushed at 48dfe1a. Each finding of the two round-1 reviews is answered below. The PR body is rewritten to match; every figure there and here is regenerated by a script named next to it, over saved logs. Correctness reviewC1 (major), fixed. Verified first: I read the kernel source, and the verifier measured on real Ubuntu archive kernels under QEMU. On 5.15.0-191 the reap's form (
C2 (minor), fixed.
C3 (minor), fixed, as the verifier proposed.
C4 (nit), fixed. C5 (nit), acknowledged, unchanged. The close paths' fd-number hazard is pre-existing and not touched here. DECISION step 5 files it with a milestone after the fix lands, as planned. On 5.10 to 5.18 the same close paths also send a cancel the kernel rejects; that part is in #682. C6 (nit), fixed. The A5 comment now says an idle SQ thread needs its NEED_WAKEUP kick, and that no tier enables SQPOLL. I changed the comment and no code, because no test could reach an SQPOLL branch. C7 (nit), disclosed. Evidence reviewE1 (major), fixed. T4 is committed as
E2 (minor), fixed. The one-worker leg now also requires each test's own summary line (
E3 (minor), fixed.
E4 (minor), fixed. The summary now says 0 keep-alive failures and 0 churn read timeouts on the switch-stall harness. The churn census (fix: reset 189, refused 16, dial reset 3) is in the body, from E5 (nit), fixed.
E6 (nit), fixed. Each of the three steps has a skipped-subtest arm, and all three fail on the indented SKIP line. The concurrency sampler now polls every 0.2 s from launch: every arm saw exactly one container, including the seccomp arms (2 samples for the unit one). E7 (nit), fixed. The body discloses the Also open from round 1m15, removed as equivalent. The verifier proved it, from code and by a counting overlay: 60,559 recv-site entries with the hold set, none changing anything, and the send site releasing 86 of 90 in the same runs. The call is removed at both recv dispatch sites; Refused dials at the revert (post hoc): still open, and not changed by this PR. The verifier read it from code:
The falsifiable hypothesis and the measurement it needs (per-revert in-process timestamps, a 1 ms LISTEN sampler, stratified by suspended epoll loops) go to PR 3, which owns the switch. Suites, lint and CI for this head are in the body. In short: 0 items go from PASS on 86d46ac to anything else on this head, across 10 whole-package cells; lint is clean in all five modules; and CI run 35449608095 had 11 of 11 jobs green on the first attempt, with its tally lines quoted in the body. |
Round 2: refused dials at a revert, measured (pre-registered)Verdict: The design was registered in
Primary: runs with a refused churn dial within ±100 ms of a revert
Stratified by whether any epoll loop was parked at ResumeAccept
Instrument
What the witness shows (post hoc, descriptive: A4)
Registered predictions
Scripts (all under
How much this rules out (added by the maintainer after review)
|
…ot reap (celeris#657) Review R1 on #681. Where the kernel rejects IORING_ASYNC_CANCEL flags, a promoted async conn under a reverse drain never left io_uring, yet its dispatch goroutine claimed the hand-off at every park: the worker's drain of the claim found the recv the feed path had armed, which only a reap can clear, and refused; the next request respawned the goroutine. A spawn and a detach-queue round trip per request, with no end while the drain lasted. asyncTransplantEligible now refuses on a worker without cancel flags, so the goroutine parks as it does with no drain set and the conn stays on io_uring (placement only). asyncCancelFlags is set before the worker starts and never written again, so the dispatch goroutine reads it race-free. Tests: TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing runs the real dispatch goroutine through three requests with the drain set and requires no claim and one goroutine for all three; on 48dfe1a it fails with 3 of 3 parks claimed, 3 goroutines and TransplantReapUnsupported = 3. Its control (control_async_park_claims_with_flags) shows the same rig sees a claim, a reap and the hand-off where the worker can reap. TestReapSuppressedAfterFailedHandOff/async_conn_retries_at_its_next_park pins the async side of a failed dup: the data that respawns the goroutine clears reapSuppressed before any park, so the hand-off is retried at the next park, once per request, and the conn leaves once the dup works.
…rejection (celeris#657) Review R2 on #681. probeAsyncCancelFlags cached a probe that never reached the kernel (a ring that could not be set up, a failed submit or wait, a missing or foreign completion) exactly as it cached -EINVAL, and New logged both at Info as "probe failed". A rejection is the kernel's answer; a probe with no answer says nothing about the kernel, and on 5.19 or later it turns the hand-off's reap off where it should work. The probe now returns asyncCancelAccepted, asyncCancelRejected (-EINVAL, or a result it does not recognise) or asyncCancelNoAnswer. Only accepted turns the reap on. logAsyncCancelProbe logs a rejection at Info and no answer at Warn when the kernel's version is 5.19 or later (Info below it), each with the reason and the kernel; the engine-selected line also carries async_cancel_probe. The probe's ring is made through newAsyncCancelProbeRing, a variable, so a test can make the probe fail before the kernel answers. The docs that said the next response is then HELD now state HOLD's condition: a worker without a provided-buffer ring. A kernel that rejects the flags has none (buffer rings arrived in 5.19 too); where one exists because the probe got no answer, a sync conn is not held and stays on io_uring. Tests: TestWorkersCarryTheAsyncCancelProbe/no_answer_is_logged_apart_from_a_rejection runs New with the probe's ring failing: on 48dfe1a plus only the ring seam it fails, the record at INFO reading "async cancel flags runtime probe failed", want WARN. TestAsyncCancelProbeClassifies gains no_answer_is_not_a_rejection and log_levels (the level matrix).
|
Round 3 is pushed at 272bcba. It answers R1 to R8 from the two round-2 re-reviews. One half of R1 is disputed, with the evidence below; everything else is fixed. The PR body is updated to match. Every figure here and there comes from a script named next to it, run over saved logs. Round 3's evidence is under Commits: 16c9719 (R1), 5e9c7f0 (R2), c15c686 (R3), 356fbe2 and 1037b66 (R4), c173711 (R5), and 1f883e2, a test-only fix for two races in round 3's own async rig, found by the local interlock run and a stress run. After those come 683c7bb (the deliberate red commit, R6) and 272bcba, its revert. The final tree is byte-identical to 1f883e2's. R1 (minor): fixed where the worker cannot reap. The
R2 (nit): fixed.
R3 (nit): fixed.
R4 (nit): fixed.
R5 (nit): fixed (docs only).
R6 (nit): done. Run 35465814564 is red on purpose.
R7 (nit): fixed. P4's figures (60,559, 6, and 86 of 90) now cite R8: fixed. Residual 5 is rewritten around the measured result: Checks on this head.
|
…er (celeris#681 N3) The zero value of asyncCancelProbe was asyncCancelAccepted, the one answer that turns the hand-off's reap on, so an answer that was never set read as ON. The zero value is now asyncCancelNoAnswer, which keeps the reap off. TestAsyncCancelProbeClassifies/zero_value_is_no_answer pins it. asyncTransplantEligible's doc now says that a probe with no answer has the same effect there as a rejection.
…its own class, and warn with the errno (celeris#681 N2) classifyAsyncCancelProbe read any completion other than an acceptance (res >= 0, -ENOENT) or -EINVAL as a rejection, and New logged it at Info, even on a 5.19+ kernel where the flags exist. It is now its own class, asyncCancelUnexpected: the reap stays off, the reason names the errno (for example "cqe.res=-9 (EBADF)"), and New logs it at Warn on every kernel. TestAsyncCancelProbeClassifies/unrecognised_answer_is_its_own_class pins the class, the errno and the Warn; the classification table and log_levels gain the new class. The docs that list what keeps the reap off name it.
…e, and wait again after a cut-short wait (celeris#681 N1) probeAsyncCancelFlagsCached kept whatever the probe returned for the life of the process, a probe that got no answer included. One transient failure, such as EMFILE or ENOMEM setting up the probe's ring while the adaptive engine builds its io_uring engine under load, kept the hand-off's reap off from then on. Now only the kernel's answer, accepted or rejected, is cached. After no answer, or an answer the probe does not recognise, the next New probes again. Calls are serialised by a mutex. The probe waits for its cancel's completion only when it is not there after the submit. SubmitAndWaitTimeout returns nil both when that wait times out and when a signal cuts it short (EINTR), so the probe used to read an interrupted wait as no answer. A wait that returns early with nothing to read is now repeated once, for the rest of the 500 ms. Test seams: runAsyncCancelProbe (the probe the cache runs), and the probe's submit, wait and wait timeout. New subtests: TestWorkersCarryTheAsyncCancelProbe/only_an_answer_is_cached and /unexpected_answer_keeps_the_reap_off_and_is_not_cached, and TestAsyncCancelProbeOnThisKernel/a_cut_short_wait_is_retried_once.
…celeris#681 N4) The branch for a kernel that rejects the cancel flags now logs what it saw (TransplantReapFailed, TransplantHeld, adopted), so a run on such a kernel shows the branch ran. Test-only.
…id not find the cancel flags (celeris#681 N4) Running the witness list on Ubuntu 5.15.0-191 under QEMU (N4) found TestHandoffHasNothingInFlight/async failing there on round 3's head and on this one alike: "0 of 128 busy conns were handed off, want all". Without the cancel flags a promoted async conn cannot be reaped, so it is never offered for the hand-off and stays on io_uring (asyncTransplantEligible, celeris#681 R1). The subtest wanted every conn handed off on every kernel. Where the engine's probe did not find the flags, the async subtest now wants none of the promoted conns handed off, with the loss checks unchanged (no client error, W1 = W2 = 0). A new case, async_without_cancel_flags, takes the probe's answer as a rejection (runAsyncCancelProbe), so that branch runs on every kernel, CI's included. Test-only.
|
Round 4 is pushed at 5ce4986. It answers N1 to N6 from the two round-3 re-reviews. The PR body is updated to match, and its first line now says the PR must be squash-merged (N6). Every figure here and there comes from a script named next to it, run over saved logs; round 4's evidence is under Commits: 30f9f70 (N3), 0c01c1c (N2), 0265e82 (N1), and two test-only ones from the QEMU runs N4 asked for: 92a8170 (the rejecting branch logs what it saw) and 5ce4986 (an engine-level test that could not pass on a kernel without the cancel flags — see N4). N1 (nit): fixed. Only the kernel's answer is cached, and a wait a signal cut short is repeated once.
N2 (nit): fixed. A completion the probe does not recognise is its own class.
N3 (nit): fixed. The zero value of N4: run on the real kernels. Round 2's rig, reused: the same Ubuntu-archive
N5: done — I stopped naming it. The mutation set no longer claims that test as a killer anywhere: not in
N6: done. The PR body now opens with: "Squash-merge this PR. Its branch carries 683c7bb, "DO NOT MERGE: plant the no-HOLD/REAP engine …", the deliberate red commit round 3's R6 asked for, reverted by 272bcba. A squash keeps that commit out of Checks on this head.
|
…e async hand-off subtest demands (celeris#681 M1) TestHandoffHasNothingInFlight/async chose between "all 128 promoted async conns handed off" and "none of them" from e.asyncCancelFlags alone. That boolean is false for every probe answer except an acceptance, so a probe that got no answer -- its private ring failing with EMFILE or ENOMEM, which celeris#681 N1 made a per-New event -- silently downgraded the subtest's strongest assertion into one any run with the reap off passes trivially. The subtest is the registered deterministic killer of mutants m01 and m04, so the mutation claim inherited the weakness. The branch is now chosen by the running kernel: from 5.19, which is where IORING_ASYNC_CANCEL flags exist, the strong branch is demanded and an engine whose reap is off fails the subtest, naming the probe's class from New's own "io_uring engine selected" record rather than probing a second time. Only a kernel that predates 5.19 may take the weak branch. async_without_cancel_flags, which forces the no-flags path on every kernel, is unchanged.
…, and why the two probes in TestReapOnTheRunningKernel stay (celeris#681 N-a, N-d) runAsyncCancelProbe, asyncCancelProbeSubmit, asyncCancelProbeWait and asyncCancelProbeTimeout are read inside probeAsyncCancelFlagsCached's asyncCancelMu critical section and replaced by tests that hold no lock. Nothing in engine/iouring calls t.Parallel(), so nothing races today; the constraint is now written where the seams are declared, with the grep that checks it, rather than enforced with a mutex-taking setter per seam. TestReapOnTheRunningKernel's two probeAsyncCancelFlagsCached calls can be two probes since N1 cached only an answer. They stay: neither subtest weakens a check on the answer it got, and probe_matches_the_kernel Fatalfs when the answer disagrees with what the kernel does with a real reap.
EngineMetrics.TransplantReapUnsupported had a 29-character fragment line mid-paragraph; asyncTransplantEligible's comment had a 131-character line among neighbours that wrap at about 78. Text unchanged.
|
Round 5 is pushed at edca5d2. It answers M1 and N-a to N-e. The PR body is updated (the round-5 paragraph and residual 7). Every figure here comes from a verdict file named beside it, each regenerated by a saved script; round 5's evidence is under Commits: 8e09151 (M1), 26f9cb9 (N-a, N-d), edca5d2 (N-b). M1 (minor): fixed. The kernel decides what the subtest demands, not the probe's answer. The rule I implemented, in
Failing first — one seam,
The two base arms are the discriminating pair: the same tree, the same test, the seam the only difference, and it turns a real check into a vacuous one. The head's failure message is
On real kernels. Round 2's QEMU rig again — the same Ubuntu-archive
That is the rule end to end: the weak branch survives exactly where the kernel genuinely predates 5.19, and a probe with no answer is fatal from 5.19 on. ( The mutation claim it carries. The subset re-run on this head is every mutant whose must-fail list names N-a (nit): documented, not guarded — and here is why. The four seams ( N-b (nit): fixed. N-c (nit): added. Residual 7 in the PR body now states round 4's per- N-d (nit): I leave it, and said so in the code. N-e (nit, pre-existing): recorded as accepted.
Checks on this head (all local runs: one container at a time,
|
…ines (celeris#662, celeris#657) The pause keeps a listener open for about 1.5 s with TCP_DEFER_ACCEPT cleared, so for that long the engine a switch is leaving keeps its SO_REUSEPORT share of NEW connections. Nothing pinned where those go, and that is exactly what failed this PR's base-vs-branch suite gate on 2026-09-18: with the linger in and celeris#657 open, an async 2048-connection ramp left 972-1170 conns on the outgoing epoll after 20 s, err=0. Not a loss -- placement. celeris#657 is now fixed (#681, #687) and this is the test that holds the joint property, so the interaction cannot regress silently. TestLingerArrivalsReachTheIncomingEngine, in engine/epoll and engine/iouring, with an async handler and the default configuration: * BeginPauseAccept, then wait for the socket witness to read the option off every listener -- the linger has begun; * PHASE A: six connections arrive, are served, and go silent while NO drain is set. Their dispatch goroutines park with nothing owed for them (epoll: askAtPark returns early with no drain; io_uring: runAsyncHandler's claim is not reached). This is the flap case, a second switch inside the first one's linger; * StartTransplant, still inside the linger; * PHASE B: six more arrive, are served, and go silent with the drain set. They join the live set at the TAIL, past the sweep cursor -- the case the cycle rule exists for (celeris#657 R2 MAJOR-1); * then every listener closes and the engine must hold none of them. It asserts what a client and an operator can each check: no dial refused and every arrival accepted by THIS engine (OnConnect is the witness), every arrival answered 200, every one adopted by the target, ActiveConnections back to 0, and the hand-off ledger balanced (io_uring also requires TransplantHandoffInFlight 0). The target serves what it adopts. A target that only holds the descriptor cannot tell a connection that was MOVED from one that was DROPPED, because either way its client never gets an answer -- the first draft of this test read 12 timeouts and 12 adoptions and could not say which had happened. Both controls lose, measured in one container each, mutation restored and sha256-verified (evidence celeris-662/r6/rebase/logs/c05..c08): * linger 0 (celeris#662 off): 12 of 12 dials refused, 0 accepted, 0 served, on both engines. * the sweep and the park-boundary ask removed (celeris#657 off): epoll adopts 0 of 12 and still holds 12; io_uring adopts 6 of 12 and still holds the 6 of phase A. io_uring's phase B moves without the sweep because its own park claims it, which is why phase A is in the test: it is the only shape on that engine that the sweep alone carries. CI: the name joins the celeris#662 linger step's tally, which runs both packages with -v, forbids every skip and asserts the exact PASS count.
…on the two ramp tests (celeris#662) The quarantine comment was written when celeris#657 was open and said the two Ramp tests fail on it. Both halves of #657 are on main now (#681, #687), so that paragraph is stale in the tree this branch produces. Measured on the rebased head, 3 containers, golang:1.27, --cpus 4, memlock 128 MiB (4 io_uring workers), seccomp unconfined, -race -count=1, one package per container (evidence celeris-662/r6/rebase/logs/s06,s07-rep1,s08-rep2, regenerated by tools/tally.sh): TestRampH1Sync PASS 3/3 TestRampH1Async PASS 3/3 every high phase, both variants: epoll=0 of 2048, settled_in 2.50-2.55 s err=0 over 6,755,764 requests; err census total=0, no class at all That is the registered prediction on celeris#687 (comment 5749130618) held: the residue that failed this PR's own G5 gate -- 972-1170 of 2048 left on the outgoing epoll after 20 s -- is what #687's sweep removes. The names stay on -skip regardless. The lift rule is >= 6 GitHub-hosted re-runs of identical bytes, the same rule #687 applied to TestBidirectionalFlapAsync, and a 2048-connection ramp on a shared runner is the shape celeris#686's dial-time resets were seen in. Lift it from a green PR run.
… resolve it: close paths, hijack, shutdown (celeris#685) Fixed files are off, so a recv or send SQE names its descriptor by number and the kernel resolves the number when it issues the op. finishClose and finishCloseDetached queued the ops' cancels and closed the descriptor at once, while an op could still be issued: a recv still in the SQ ring (a promoted connection's re-arm), or one linked behind a SEND that the kernel issues only after the SEND completes. A sibling worker given the freed number before this worker's next enter then had its new connection's request read by the old recv, dropped as stale_recv_data_closed, and the client was never answered (celeris#715: 80 of 80 in #781's trial). The rule, as #681 applied it to the hand-off: a close path does not release the number while the kernel owes an op on it (fdOwed: kernelInflight > 0, which counts every recv and send from the moment its SQE is written). It does everything else as before and hands the descriptor to its pendingRelease entry; drainPendingRelease closes it where it releases the connState, at the last owed op's terminal CQE. The socket's read side is shut down too (SHUT_RD on the H1 fast path, SHUT_RDWR where the path already half-closed), so an owed recv ends as soon as it is issued, even a linked one no cancel can find, and even where cancels fail (#682). No io_uring_enter is added; the close(2) moves one iteration later. The worker does not park while such a close is outstanding. hijackConn keeps the socket open under the hijacker, so it submits its cancels before handing the socket over when an op is owed (a multishot recv stays armed across its request). Worker shutdown ends the ops owed on connection descriptors (cancel, SHUT_RDWR, run the ring until their terminal CQEs, bounded at 250 ms) before it closes any, and before shutdownDrivers, which relies on nothing being submitted after it. EngineMetrics gains CloseFDDeferred (a rate) and CloseFDForced (the backstop closing a descriptor with an op still owed; must stay 0). The recv-theft CI job gets a detector control: with fdOwed forced false the three judging trials must fail every run. Fixes #685 Refs #715
… resolve it: close paths, hijack, shutdown (celeris#685) Fixed files are off, so a recv or send SQE names its descriptor by number and the kernel resolves the number when it issues the op. finishClose and finishCloseDetached queued the ops' cancels and closed the descriptor at once, while an op could still be issued: a recv still in the SQ ring (a promoted connection's re-arm), or one linked behind a SEND that the kernel issues only after the SEND completes. A sibling worker given the freed number before this worker's next enter then had its new connection's request read by the old recv, dropped as stale_recv_data_closed, and the client was never answered (celeris#715: 80 of 80 in #781's trial). The rule, as #681 applied it to the hand-off: a close path does not release the number while the kernel owes an op on it (fdOwed: kernelInflight > 0, which counts every recv and send from the moment its SQE is written). It does everything else as before and hands the descriptor to its pendingRelease entry; drainPendingRelease closes it where it releases the connState, at the last owed op's terminal CQE. The socket's read side is shut down too (SHUT_RD on the H1 fast path, SHUT_RDWR where the path already half-closed), so an owed recv ends as soon as it is issued, even a linked one no cancel can find, and even where cancels fail (#682). No io_uring_enter is added; the close(2) moves one iteration later. The worker does not park while such a close is outstanding. hijackConn keeps the socket open under the hijacker, so it submits its cancels before handing the socket over when an op is owed (a multishot recv stays armed across its request). Worker shutdown ends the ops owed on connection descriptors (cancel, SHUT_RDWR, run the ring until their terminal CQEs, bounded at 250 ms) before it closes any, and before shutdownDrivers, which relies on nothing being submitted after it. EngineMetrics gains CloseFDDeferred (a rate) and CloseFDForced (the backstop closing a descriptor with an op still owed; must stay 0). The recv-theft CI job gets a detector control: with fdOwed forced false the three judging trials must fail every run. Fixes #685 Refs #715
Summary
This is PR 2 of 3 for #657. It fixes face 2, the silent request loss at the io_uring → epoll hand-off, with one rule: a connection leaves io_uring only when no read can still resolve its descriptor.
On
main, a revert hands a keep-alive connection to epoll while its next RECV is still armed, either linked behind the response's SEND or standalone. Closing the original fd does not end that RECV. It later reads the client's next request off the socket epoll now owns, or, once the fd number is reused, another connection's request. Its CQE carries the old (fd, generation), so it is dropped as stale and the request is gone, while the #624 hand-off ledger keeps balancing. PR 1 (#676) made this countable:StaleRecvData{Transplanted,Unattributed}, a request read by a recv that outlived its hand-off.TransplantHandoffInFlight, a hand-off made with an op in flight.This PR takes both to 0. In the integration campaign (run on round 1's fix, 308ffca):
main, now adaptive: a revert can refuse dials — epoll ResumeAccept returns before any loop listens, then io_uring closes its listener (3.4% of reverts) #683 (residual 5).Round 2 (48dfe1a) answered the review of round 1: the reap is safe on kernels without async-cancel flags and after a failed dup, the double-claim counter is split, T4 is committed and gated in CI, and the one-worker leg requires zero lost requests.
Round 3 answers the round-2 re-reviews:
CELERIS_REQUIRE_UPSWITCH=1; the engine-level tests assert the must-stay-0 counters (R4).Round 4 (this head) answers the round-3 re-reviews:
-race(N4): the rejecting branch ofprobe_matches_the_kernelran and passed on 5.15 for the first time. That run also found an engine-level test asking for a hand-off that a kernel without the flags never makes; it is fixed here.Round 5 (this head) answers the round-4 re-review:
TestHandoffHasNothingInFlight/asyncno longer picks its assertion from the probe's answer alone. From Linux 5.19, the release that added the flags, it demands the strong branch — all 128 promoted async connections handed off — and fails, naming the probe's class fromNew's own record, where the engine's reap is off. Only a kernel that predates 5.19 may take the weak branch, andasync_without_cancel_flags, which forces the no-flags path on every kernel, is unchanged (M1). Before this, a probe that got no answer on a 6.x runner — which N1 made a per-Newevent — silently turned that assertion into one any run with the reap off passes trivially, and the subtest is the registered deterministic killer of mutants m01 and m04.TestReapOnTheRunningKernel's two probes stay, because neither of its subtests weakens a check andprobe_matches_the_kernelfails loudly on a disagreement (N-d); two comments round 4 left ragged are reflowed (N-b); and the residual below is round 4's new per-Newcost (N-c).Refs #657. PR 3 closes it.
What changes, and why each part exists
Everything is in
engine/iouring:fd_lifetime.go;transplant_source.go,worker.go,cqe.go,conn.go,handoff_loss.go,probe.goandengine.go.Eight new
EngineMetricsfields are added inengine/engine.goand summed over both sub-engines inadaptive/engine.go.P1. The R0 gate at both hand-off sites.
tryTransplant(sync) andfinishAsyncTransplant(promoted async) refuse whilerecvArmed || kernelInflight != 0 || zcNotifPending. When the only op in flight is the recv, they start a REAP (P2). The dup, detach, close and adopt sequence the two sites shared is now one function,handOff, reached only after the gates.Why: this is the precondition of every measured loss. On the base,
TransplantHandoffInFlightequalled the number of hand-offs in every one-worker run.P2. REAP, a reported cancel of exactly that recv.
prepCancelUserDataReportedmatches on the recv's user_data, generation included, so it cannot touch the next owner of the fd number. Its own tag isudTransplantReap = 0x09<<56.-ECANCELED, whichhandleRecvroutes first (reapOutcome), before the genericres <= 0close branch.-ENOENT) is retried on the next loop iteration while the recv is still armed. A miss is never followed by a hand-off.TransplantReapFailed, is not retried, and is never followed by a hand-off. (Round 2.)IORING_ASYNC_CANCEL_ALL, which exists from Linux 5.19; through 5.18 the kernel returns-EINVALand leaves the recv armed.probeAsyncCancelFlags, submits exactly the reap's SQE against a user_data nothing carries. Round 3 (R2), round 4 (N2, N3): it answers one of four things. Accepted (res >= 0or-ENOENT) turns the reap on. Rejected (-EINVAL) is the kernel's answer, and is what every kernel before 5.19 gives. Unexpected is a completion the probe does not recognise; its reason names the errno ("cqe.res=-9 (EBADF)"). No answer (the ring could not be set up, the submit or the wait failed, no completion, a foreign completion) says nothing about the kernel. The last three keep the reap off, and the zero value of the answer is "no answer", so an answer never set cannot turn it on.Newlogs a rejection at Info, an unexpected answer at Warn on every kernel, and no answer at Warn when the kernel's version is 5.19 or later (Info below), each with the reason and the kernel. The engine-selected line carriesasync_cancel_flagsandasync_cancel_probe.createWorkerscopies the answer to every worker.Newprobes again and logs again, so one transientNewRingfailure — EMFILE or ENOMEM while the adaptive engine builds its io_uring engine under load — no longer keeps the reap off for the life of the process; a mutex serialises the probes. The probe's wait for its cancel's completion is repeated once when it comes back early with nothing to read:SubmitAndWaitTimeoutreturns nil for a wait a signal cut short (EINTR) as well as for its timeout, so only an early return can be the former.TransplantReapUnsupported. A connection the worker serves itself keeps its recv until that recv completes on its own, and its next response is then HELD, where the worker has no provided-buffer ring (HOLD needs none; a kernel that rejects the flags has none, since buffer rings also arrived in 5.19). With one it stays on io_uring.asyncTransplantEligiblerefuses, so its dispatch goroutine parks as it does with no drain set, and it stays on io_uring (placement only). Before, the goroutine claimed at every park, the claim was refused, and the next request respawned it: a spawn and a detach-queue round trip per request for as long as the drain lasted.SetNonblocksetsreapSuppressed: no reap is placed for that connection until it next receives data (round 2). The dup goes through aWorker.dupFDtest seam. Round 3 (R5): only data and the release clear the flag, not the end of the drain, and the field's doc now says what that does across drains.reapOutcomeconsumes the reaped recv's last completion, so it does whathandleRecvdoes at any recv's last completion: a flagged provided buffer goes back to the ring andrecvLinkedis cleared (round 2, defensive).Why a count and not a bool: this is the #484/#596 lesson. A stale miss must not clear the state a newer reap needs; mutant m05 shows it. Why its own tag: so a WebSocket pause's cancel (
udRecvCancel, #596) is never taken for a reap, or the reverse.P3. HOLD. While a drain is set and the worker has no provided-buffer ring, a connection the hand-off would accept has its response flushed unlinked, with no recv armed at all. Its SEND completion then finds nothing in flight and hands the connection off; the client's next request waits in the socket buffer for epoll. Both response tails go through one helper,
respondAndArm. With no drain set, the SQEs are exactly those placed before:TestNoDrainSQESequenceIsUnchangedpins this, and it passes on the base too.P4. The release guarantee.
releaseHoldruns, unconditionally, after the transplant attempt at both SEND dispatch sites: the inlined one andprocessCQE's, which the listener-close harvest uses.checkTimeoutsrescues any held connection a path left unreleased and counts it inTransplantHoldRescued. This must stay 0.releaseHoldSlowreturns at its egress check. Measured on 86d46ac with a counting overlay, as tallied by the verifier'sverify/tools/v4_tally.pyoververify/v4/runs(verify/v4/V4-TALLY.txt; re-run in round 3 with identical output, itspr/V4-TALLY-regenerated.txt): the inline recv site was entered 60,559 times with the hold set and changed nothing, andprocessCQE's recv site 6 times, also nothing. The send site released 86 of its 90 entries in the same runs (the positive control).P5. One owner per hand-off (A6, a pre-existing bug).
tryTransplantleaves a promoted async connection alone when its dispatch goroutine has claimed its own hand-off (transplantPending, read underasyncInMu).finishAsyncTransplantacts only for the connState that still owns its slot. The two refusals are two counters (round 2):TransplantClaimDeferred(a rate) is the first. It is ordering, not a double claim: a completion of the connection landed between the goroutine's park and the drain of its claim, and it is counted before any other gate.TransplantDoubleClaimcounts only a connState that no longer owns its slot. In async mode, only another hand-off of the same connection vacates the slot that way: a close marks the queued claim first, and a hijack is refused. So it must stay 0, and a release gate of 0 is meaningful.P6. Submit pending SQEs before a worker parks (A5). When the worker is about to park indefinitely (no listener, no connections), it submits what the iteration queued. Pending close-path cancels were measured sitting unsubmitted at parks (1 to 24 of them) until something woke the worker. The comment says an idle SQ thread needs its NEED_WAKEUP kick; no tier enables SQPOLL today.
EngineMetricsfields (io_uring-only, cumulative, zero elsewhere; the adaptive engine sums both sub-engines):TransplantHeldTransplantReapsTransplantReapMissesTransplantReapFailedTransplantReapUnsupportedTransplantHoldRescuedTransplantDoubleClaimTransplantClaimDeferredThe gate this PR supports: W1 == 0 and W2 == 0, plus
TransplantDoubleClaim,TransplantHoldRescuedandTransplantReapFailed== 0. The rates are reported, not gated. Since round 3, T4 and the engine-level tests assert exactly this gate.Tests
engine/iouring/fd_lifetime_test.go, synthetic completions through a ring-backed worker:TestTransplantNeverHandsOffArmedRecv,TestTransplantReapMissIsRetried,TestHoldReleasedWhenDrainStops,TestHoldRescuedByCheckTimeouts,TestOneOwnerPerHandoffandTestNoDrainSQESequenceIsUnchanged;TestTransplantReapFailureIsNotRetried,TestNoReapWithoutAsyncCancelFlags,TestReapSuppressedAfterFailedHandOffandTestReapedRecvLeavesNoLinkOrBuffer;asyncParkFixture, which records the goroutine each request ran on and sees a park insync.Cond.Wait):TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing(R1): three requests with the drain set on a worker without cancel flags; no claim, one goroutine for all three, nothing placed but the feed path's RECVs. Its control,control_async_park_claims_with_flags, shows the rig sees the claim, the reap and the hand-off where the worker can reap.TestReapSuppressedAfterFailedHandOff/async_conn_retries_at_its_next_park: the async side of a failed dup (see R1 in the round-3 reply).reapSuppressedis never set at a park; the hand-off is retried once per request while the dup fails and the connection leaves once it works.engine/iouring/async_cancel_probe_test.goandasync_cancel_probe_log_test.go:TestAsyncCancelProbeClassifies, withno_answer_is_not_a_rejectionandlog_levels(round 3), and round 4'szero_value_is_no_answer(N3) andunrecognised_answer_is_its_own_class(N2: the class, the errno in the reason, and the Warn on 5.15 and 6.8, for-EBADF,-ECANCELEDand-EPERM);TestAsyncCancelProbeOnThisKernel: logs the probe's answer beside the kernel version (round 3: no longer a verdict), drives the rejection path with a flag bit no kernel defines, and (round 4)a_cut_short_wait_is_retried_once, which holds the cancel back at the submit and cuts the wait short once, twice, and once at full length;TestWorkersCarryTheAsyncCancelProbe, withno_answer_is_logged_apart_from_a_rejection(round 3:Newwith the probe's ring failing), and round 4'sonly_an_answer_is_cached(three calls per class of answer, counting the probes, then twoNews across a no-answer) andunexpected_answer_keeps_the_reap_off_and_is_not_cached;TestReapOnTheRunningKernel, where the real kernel completes every op, and (round 3)probe_matches_the_kernel, which forces the flags on, lets the running kernel answer the real reap, and requires the probe's answer to be what the kernel did.engine/iouring/fd_lifetime_engine_test.go, a real engine under keep-alive load:TestHandoffHasNothingInFlight,TestStaleRecvDataCounted(the join as a test: client errors == W1),TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappenandTestWorkerParksWithNothingPending. Round 3: every load run fails on a non-zeroTransplantDoubleClaim,TransplantReapFailedorTransplantHoldRescued. Round 4 (N4): where the engine's probe did not find the cancel flags,TestHandoffHasNothingInFlight/asyncwants none of the promoted connections handed off — such a connection is never offered for the hand-off, so it stays on io_uring — with the loss checks unchanged; it used to want all of them on every kernel, and so could not pass on 5.15.0-191. A new case,async_without_cancel_flags, takes the probe's answer as a rejection, so that branch runs on every kernel, the runner's included.adaptive/flap_conns_per_ring_test.go:TestFlapConnsPerRingis T4. It uses two workers per engine, so its 256 keep-alive connections are 128 per io_uring ring, and runs three promote/revert cycles under load. It requires 0 client errors, and since round 3 (R4), from the adaptive engine'sMetrics(): W1 = 0, W2 = 0,TransplantDoubleClaim,TransplantHoldRescuedandTransplantReapFailed= 0, and at every switch that the engine switched away from detached at least all 256 and holds none while the one switched to adopted at least all 256 and holds all of them. Its memlock, io_uring-availability andadaptive.Newskips all fail instead underCELERIS_REQUIRE_UPSWITCH=1.Failing first, round 3. Source: round 3's
ff/VERDICT.txt, from itstools/ff_verdict_r3.pyovertools/run_ff_r3.sh's logs.TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing(R1), on 48dfe1a plus only this head's two fd-lifetime test files: "the dispatch goroutine claimed the hand-off at 3 of 3 parks, and the 3 requests ran on 3 goroutines, want 0 claims and 1 goroutine", and "TransplantReapUnsupported = 3, want 0". It is the only failing subtest there; its control and the async suppression pin pass on 48dfe1a too.TestWorkersCarryTheAsyncCancelProbe/no_answer_is_logged_apart_from_a_rejection(R2), on 48dfe1a plus only the probe-ring seam (round 3'sff/ovl/FF-48-R2SEAM/probe.go.seam.diff): "a probe that got no answer on kernel 7.0 was logged at INFO as "async cancel flags runtime probe failed: …", want WARN". The other R2 subtests use the new result type, which 48dfe1a does not have.CELERIS_REQUIRE_UPSWITCH=1: 48dfe1a's SKIPs ("io_uring unavailable: needs both sub-engines"), and this head's FAILs ("… -- CELERIS_REQUIRE_UPSWITCH=1 forbids skipping").ff/VERDICT.txt).Failing first, round 4. Source:
ff/VERDICT.txt, fromtools/ff_verdict_r4.pyovertools/run_ff_r4.sh's logs. Each of round 4's four checks is run on 272bcba plus only round 4's test seams (runAsyncCancelProbe, and the probe's submit, wait and wait timeout:ff/ovl/FF-272-SEAM/probe.go.seam.diff), under the subtest name this head gives it:zero_value_is_no_answer(N3): "the zero value of asyncCancelProbe reads as accepted: an answer that was never set would turn the hand-off's reap on".unrecognised_answer_is_its_own_class(N2): "classifyAsyncCancelProbe(-9) = rejected (…), want a class of its own, "unexpected"", and the record logged at INFO where WARN is wanted.only_an_answer_is_cached(N1): "three calls with the probe answering no answer ran it 1 time(s), want 3", and "the first New's probe got no answer, the second's is accepted: asyncCancelFlags false then false, probes run 1; want false, true and 2".a_cut_short_wait_is_retried_once(N1): "wait cut short once: the probe gave (no answer, "no CQE produced for the cancel (waited 500ms)") after 1 wait(s), want accepted after 2".FF-272-SEAMCTL), so the seams change no behaviour, and the full-length-wait check — which is not a behaviour change — does not fire there.CI
ci.ymlis unchanged since round 2. The new tests are subtests of the 29 tests the witness step names, so its tally and its ban on any SKIP line (indented ones too) cover them.unitjob, witness step. It names 29 tests: PR 1's 11 and PR 2's 18. It runs with-v,CELERIS_REQUIRE_IOURING_WORKERS=1and an exact tally of top-level=== RUNand--- PASS: Name (lines, and it fails on any--- SKIPline. Everyceleris657 engine workers=Nline must read 1.adaptivejob, race step (memlock raised). It tees its log and tallies T4: exactly one top-level RUN and one PASS, no SKIP line for T4, and exactly oneS0T4 RESULTline readingconns=256 wE=2 wI=2 conns_per_ring=128 … err=0. The RESULT line now also carriesw1= w2= doubleclaim= holdrescued= reapfailed= held= reaps=, aftererr=, so the tally's pattern is unchanged.adaptivejob, one-worker leg (runner memlock). Exactly 2 RUN and 2 PASS, no SKIP line, everyio_uring engine listeningline atworkers=1, and each test's own summary line aterr=0. The verdicts tolerate up to 64 lost requests; a base flap run PASSED with 56.Departures and what they imply.
Resources.Workers: 1variants of the revert tests.resource.MinWorkersis 2 (resource/config.go:147-148), so they cannot be built.minMemlockPerWorker), so it starts one.workers=1. That is intended: the shape must then be pinned another way.Every tally, proven both ways. Round 2 ran 28 arms at 48dfe1a from the verbatim
run:blocks (bash --noprofile --norc -eo pipefail, asudoshim, onegolang:1.27container per arm) and all 28 proved their side (round 2'sinterlock/48dfe1a/VERDICT.txt,tools/il_verdict_r2.py). Withci.ymlunchanged, round 3 re-ran the arms whose run changed. Source: round 3'sinterlock/1f883e2/VERDICT.txt, from itstools/il_verdict_r3.py:wE=2 wI=2 … err=0; package 76 PASS, 0 FAIL, 0 SKIPgo testfails. In round 2 T4 skipped,go testwas ok, and the tally refused the SKIP.wE=3 wI=3go testok, err=1588 of 8 arms prove their side, each in its own container, one at a time.
Round 4 re-ran the three GOOD arms on this head, from the same verbatim
run:blocks, one container each (interlock/5ce4986/VERDICT.txt,tools/il_verdict_r4.py).ci.ymlis byte-identical to 272bcba's (git diff 272bcba HEAD -- .github/is empty), so the other arms are round 3's and stand.async_without_cancel_flagscase builds one).cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=2109529 err=0, package 0 FAIL, 0 SKIP.On the runner, both ways (R6). Run 35465814564 was red on purpose. Its commit, 683c7bb ("DO NOT MERGE"), planted the no-HOLD/REAP engine and turned T4's and the revert tests' verdicts into logs, so only the steps' own tallies could fail. From its real job logs (round 3's
ci/683c7bb/SUMMARY.txt):go testpassed (76 PASS). Its T4 tally refused the RESULT line (err=188 w1=188 w2=398, "at two workers with err=0: 0").Commit 272bcba reverts it; its tree is identical to 1f883e2's. Run 35465988505 on it is green. Both runs are in round 3's
RED-RUNS.txt.How it was checked
mainb79888e. That is round 1's base (985a386 + feat(engine): count the requests an io_uring hand-off drops, and the hand-offs made with an op in flight (celeris#657) #676) plus docs: state the TLS posture, document the tuning environment variables, and give overload a doc.go #677, test(wakefd): make TestConcurrentSetAndSignal's post-Set signal deterministic #678 and test(engine): wait for every client before counting accepts in the #662 queued-accept rig #680, all test-only or docs.golang:1.27,--cpus 4 --cpuset-cpus 0-3, one container per observation and one at a time, linux/arm64, kernel 7.0.12-linuxkit. The lint containers use their own images, also one at a time.--- PASS/FAIL/SKIP: Name (lines are tallied. A SKIP is not a pass.workers=is checked in every io_uring run, and a run with the wrong count is void.evidence/celeris-657/impl/:pr2-r4/, which is the base of every path given here without a round;pr2-r3/;pr2-r2/fix/, the verifier's inpr2-r2/verify/and the refused-dial measurement inpr2-r2/measure/;pr2/(for exampleship/).Mutation (round 4). Tool:
tools/mkmut_r4.py; runner:tools/run_mut_r4.sh; verdict:mut/VERDICT.txt, fromtools/analyze_mut_r4.py; the re-scoped mutants are explained inmut/VERDICT-NOTE.txtandmut/EXPOSURE.txt(tools/exposure_census_r4.py).CELERIS_REQUIRE_IOURING_WORKERS=1(75 mutants);CELERIS_REQUIRE_UPSWITCH=1(6).New; the 5.19 threshold off by one; the ring seam ignored; a probe that always says rejected; and each counter that must stay 0 planted against the engine-level tests and against T4, plus the reverse drain turned off.workers=1for all 8 engines, and every T4 run's RESULT line readwE=2 wI=2. 0 survivors.TestStaleRecvDataCountedreaches the branch m04 changes only if a connection's recv happens to be armed when the drain examines it. Across the runs of rounds 2 to 4 that exposure ranges from 0 to 88 of 128 and is 0 in 6 of 18 (mut/EXPOSURE.txt,tools/exposure_census_r4.py), so its FAIL was never guaranteed.TestTransplantNeverHandsOffArmedRecv, which drives the branch with synthetic completions, andTestHandoffHasNothingInFlight/async, where every connection is promoted and a promoted connection always leaves through the reap (W2 = 128 in both m04 runs).TestTransplantNeverHandsOffArmedRecvandTestTransplantHandoffInFlightCounts{TryTransplant,FinishAsyncTransplant}("the R0 gate is gone").TestStaleRecvDataCountedstill runs in every one of those mutants'-runsets, andmut/VERDICT.txtreports what it did; it is simply not required.Whole-package suites, this head against 272bcba. One container per tree per cell, with
-race -v, count=1. Source:suites/COMPARE.txt, fromtools/compare_suites_r4.pyovertools/run_suites_r4.sh's logs../engine/iouring, 8 MiB, as the unit job runs it./engine/iouring, 128 MiB,CELERIS_REQUIRE_IOURING_WORKERS=1./adaptivewith the CI-skip, 8 MiB./adaptive, 128 MiB,CELERIS_REQUIRE_UPSWITCH=1WARNING: DATA RACEor a panic, and every 8 MiB run's io_uring engines are at one worker../engine/iouring, absent on 272bcba and PASS on this head in both memlock cells:zero_value_is_no_answer,unrecognised_answer_is_its_own_class,only_an_answer_is_cached,unexpected_answer_keeps_the_reap_off_and_is_not_cached,a_cut_short_wait_is_retried_onceandTestHandoffHasNothingInFlight/async_without_cancel_flags.tools/t4_switch_census_r4.py→t4-switch-census.txtcovers this round's unmutated T4 runs (the interlock T-GOOD arm, the mutation up128 control and the 128 MiB adaptive suite cell). All 18 switches moved and adopted 256 connections, and all 3 RESULT lines readerr=0 w1=0 w2=0 doubleclaim=0 holdrescued=0 reapfailed=0.Lint, as the
lintjob runs it. golangci-lint v2.13, with the repo's config, is clean in all five modules on linux/amd64 and linux/arm64.mage -compileandmage CheckReleasepass. actionlint 1.7.12 with shellcheck is clean, and so is zizmor 1.30.0 (offline). Each linter flags a defect planted in a copy of the tree (lint/,tools/run_lint.sh; re-run on this head). gofmt lists one file,test/benchcmp_ws/bench_test.go, which has been untouched since #262 and is listed on 272bcba too; no CI step runs gofmt.GitHub CI on this head. Runs 35474330959 (CI) and 35474327611 (this PR's analysis) on 5ce4986: ubuntu-24.04, linux/amd64, kernel 6.17.0-1022-azure, go1.27.0. All 11 jobs were green at the first attempt, and none was re-run. From the real job logs (
ci/5ce4986/SUMMARY.txt,tools/summarize_ci_r4.py, which counts worker witnesses only from engine and test log lines):want 29, ran 29, passed 29, SKIP lines 0, engines 8 at workers=1: 8, with 44 subtests PASS (round 3: 38; round 4's six are the difference). The probe answered accepted on the runner's kernel and the forced reap ran there (probe_matches_the_kernel), andasync_without_cancel_flagsPASSed, having driven the no-flags branch on that kernel: "0 of 128 promoted async conns handed off (promotions 128)".celeris#657 T4: want 1, ran 1, passed 1, SKIP lines 0, RESULT lines 1, at two workers with err=0: 1. The RESULT line readscycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=616472 err=0 w1=0 w2=0 doubleclaim=0 holdrescued=0 reapfailed=0 held=750 reaps=140, and every one of the six switches moved and adopted 256 connections.want 2, ran 2, passed 2, SKIP lines 0, io_uring engines 2 at workers=1: 2, summaries 2 with err=0: 2.The witness tests on real Ubuntu kernels under QEMU, at this head (N4). Round 2's rig, the same Ubuntu-archive
.debs —tools/qemu_inner_r4.shcompares their hashes with round 2's verifier before booting and refuses if they differ (they are identical) —qemu-system-aarch64(TCG), and the head's./engine/iouringtest binary built with-race. Source:qemu/VERDICT.txt, fromtools/qemu_verdict_r4.py;tools/qemu_plans_r4.sh,qemu_build_r4.sh,qemu_run_r4.sh+qemu_inner_r4.sh.TestReapOnTheRunningKernel/probe_matches_the_kernelcqe.res=-22 (EINVAL)TransplantReapFailed=1) and the connection left after its held response (TransplantHeld=1, adopted 1)WARNING: DATA RACE. This is the first time the rejecting branch has run anywhere.TestHandoffHasNothingInFlight/asyncfailing on 5.15 — "0 of 128 busy conns were handed off, want all" — on 272bcba and on the first round-4 head alike, because a promoted async connection is never offered for the hand-off where the flags are missing. 5ce4986 fixes the expectation; the subtest passes on 5.15 now, andasync_without_cancel_flagsruns that branch on every kernel.qemu/diag/(tools/qemu_build_diag_r4.sh,qemu_run_diag_r4.sh), which runs the same two halves for 272bcba's-racebinary, the head without-race, and the head with a diagnostic 10 s client read deadline (never committed).TestReapedRecvLeavesNoLinkOrBuffer/provided_buffer_goes_backon 5.15 and 6.8: "register pbuf_ring: invalid argument", no provided-buffer ring in those boots. 272bcba fails the same subtest on the same two kernels.read_timeouterrors on 5.15 and 5.19: TCG emulates the machine in software. Without-raceall four pass on 5.15, and with the 10 s client read deadline they pass on 5.15 and 5.19; 272bcba fails them there too. On 6.8 all four pass as they are.-race.The campaign (round 1, on 308ffca), unchanged. n = 8 runs per arm. W1 is
StaleRecvDataTransplanted + Unattributed, and p is Fisher's exact test, two-sided, on runs that FAIL. Every figure comes fromship/tools/regate.pyover the raw campaign logs (ship/regate/REGATE.txt), which imports nothing from the campaign's own analyzer.TestReverseTransplant, 1 io_uring worker, 64 connsTestBidirectionalFlap, 1 io_uring worker, 64 connsTestFlapConnsPerRing: 2 io_uring workers, 256 conns, 3 cycles)TestSwitchDiag662: 1 io_uring worker, 4 switches per run), n = 60verify/tools/v7_e5_w1_join.py; no campaign tool computed it.verify/tools/v7_e4_churn_census.py):TestReverseTransplantFAIL 5/8 andTestBidirectionalFlapFAIL 7/8, with lost == W1 in every run;TestOneOwnerPerHandoffFAIL 6/6;TestWorkerParksWithNothingPendingFAIL 6/6.WARNING: DATA RACEor apanic:line (verify/tools/v7_e5_logcount.sh).tools/validity_r2.pyreads both worker witnesses, andtools/e3_rerun.pyre-ran both rules over all 308 phase B, C and S logs (tooling/E3-RERUN.txt): 0 verdicts change, the 26 unit arms that build an engine all readworkers=1, and regate's worker check finds 0 violations.Interactions with other open issues
reapOutcometakes a recv's-ECANCELEDonly when no pause is set and no pause cancel is pending. Detached and paused connections are never held or reaped. TheTestBackpressure*suites are in the table above.Tracked elsewhere
-EINVALtoo: the close paths,hijackConn, the WebSocket pause,PauseAcceptand the driver unregister. Pre-existing and not changed here. This PR stops its own site from spinning and keeps promoted async connections parked there.ResumeAcceptreturns before any loop listens, and then io_uring closes its listener. Measured in round 2 (the measurement comment below) as pre-existing and not a regression of this PR. PR 3 owns the switch../adaptivetests skip and the adaptive job passes even underCELERIS_REQUIRE_UPSWITCH=1. T4 no longer skips there (round 3 makes it fail), which is why the race step goes red; the rest is that issue.What this PR does not do
This is PR 2 of 3:
StopTransplantbeforeResumeAccept(B0, P11). PR 3 closes adaptive: forced switches leave keep-alive connections behind — async conns miss the 1.2 s flap window in every run, and a one-worker io_uring strands sync conns until their requests fail #657.adaptivejob's-skip.Known residuals
CELERIS_ADAPTIVE_START=iouring. The rest of that kernel range's cancel breakage is io_uring: every async cancel fails with -EINVAL on Linux 5.10-5.18 (IORING_ASYNC_CANCEL flags are 5.19+) #682.-EALREADY'd cancel. The effect is placement only: a later miss queues no retry, so the connection stays until its client sends again, and then HOLD hands it off.CELERIS_IOURING_MULTISHOT_RECV=1, opt-in) disables HOLD, so only REAP applies there. No test exercises it.measure/00-PREREG.md, 104 runs per arm, the verdict computed bymeasure/tools/analyze.pyintoresults.json; the comment fix(iouring): hand a connection to epoll only when no read can still resolve its fd (celeris#657) #681 (comment)) found no regression:ResumeAccept, which returns before any loop listens, and io_uring's close of its listener; the control that makesResumeAcceptwait for an epoll listener had 0 refused dials in 80 reverts.The race exists on
main, the PR touches neither end of it, and it is filed as adaptive: a revert can refuse dials — epoll ResumeAccept returns before any loop listens, then io_uring closes its listener (3.4% of reverts) #683 for PR 3, which owns the switch.New. Round 4 (N1) caches only the kernel's answer, accepted or rejected, so a kernel that never answers — or answers with something the probe does not recognise — is probed again at everyiouring.New, serialised on a process-global mutex, with a log record each time (WARN from 5.19, INFO below it). Each probe opens and closes a private ring; in the remote case where the cancel produces no CQE at all, it holds that mutex for up to 500 ms (a wait to a 500 ms deadline, repeated once if a signal cuts it short). The price buys back the failure mode it replaced: cached, one EMFILE or ENOMEM in that private ring kept the hand-off's reap off for the life of the process. No kernel in this PR's evidence takes the path — on every one of them the first probe answers, is cached, and the cost is one probe per process — and from round 5 a 5.19-or-later kernel where it happens failsTestHandoffHasNothingInFlight/asyncrather than quietly weakening it (M1).