Skip to content

fix(iouring): hand a connection to epoll only when no read can still resolve its fd (celeris#657) - #681

Merged
FumingPower3925 merged 36 commits into
mainfrom
fix/celeris-657-fd-lifetime
Sep 20, 2026
Merged

FumingPower3925 merged 36 commits into
mainfrom
fix/celeris-657-fd-lifetime

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Squash-merge this PR. Its branch carries 683c7bb, "DO NOT MERGE: plant the no-HOLD/REAP engine …", the deliberate red commit round 3's R6 asked for, reverted by 272bcba. A squash keeps that commit out of main's history; a merge commit or a rebase would carry it in (N6).

Summary

This is PR 2 of 3 for #657. It fixes face 2, the silent request loss at the io_uring → epoll hand-off, with one rule: a connection leaves io_uring only when no read can still resolve its descriptor.

On main, a revert hands a keep-alive connection to epoll while its next RECV is still armed, either linked behind the response's SEND or standalone. Closing the original fd does not end that RECV. It later reads the client's next request off the socket epoll now owns, or, once the fd number is reused, another connection's request. Its CQE carries the old (fd, generation), so it is dropped as stale and the request is gone, while the #624 hand-off ledger keeps balancing. PR 1 (#676) made this countable:

  • W1 is StaleRecvData{Transplanted,Unattributed}, a request read by a recv that outlived its hand-off.
  • W2 is TransplantHandoffInFlight, a hand-off made with an op in flight.

This PR takes both to 0. In the integration campaign (run on round 1's fix, 308ffca):

Round 2 (48dfe1a) answered the review of round 1: the reap is safe on kernels without async-cancel flags and after a failed dup, the double-claim counter is split, T4 is committed and gated in CI, and the one-worker leg requires zero lost requests.

Round 3 answers the round-2 re-reviews:

  • On a worker that cannot reap, a promoted async connection's dispatch goroutine no longer claims the hand-off at every park (R1).
  • The cancel-flags probe tells a probe that got no answer from a kernel that rejected the flags, and warns on a 5.19+ kernel (R2); the probe tests compare the probe with what the running kernel does, not with its version (R3).
  • T4 asserts the loss witnesses, the counters that must stay 0 and every switch's move, and fails instead of skipping under CELERIS_REQUIRE_UPSWITCH=1; the engine-level tests assert the must-stay-0 counters (R4).
  • Two docs are corrected (R5). CI proved on the runner that it goes red on the no-HOLD/REAP engine (R6).

Round 4 (this head) answers the round-3 re-reviews:

  • The cancel-flags probe's answer is cached only once the kernel has given one, and a wait a signal cut short is repeated once (N1).
  • A completion the probe does not recognise is a class of its own, warned with its errno, and never turns the reap on (N2).
  • The zero value of that answer is now "no answer", so an answer never set keeps the reap off (N3).
  • The witness tests ran on Ubuntu 5.15.0-191, 5.19.0-50 and 6.8.0-138 under QEMU, with -race (N4): the rejecting branch of probe_matches_the_kernel ran and passed on 5.15 for the first time. That run also found an engine-level test asking for a hand-off that a kernel without the flags never makes; it is fixed here.
  • The mutation set no longer names a test as a killer where its exposure to the mutated branch is scheduling-dependent (N5).

Round 5 (this head) answers the round-4 re-review:

  • TestHandoffHasNothingInFlight/async no longer picks its assertion from the probe's answer alone. From Linux 5.19, the release that added the flags, it demands the strong branch — all 128 promoted async connections handed off — and fails, naming the probe's class from New's own record, where the engine's reap is off. Only a kernel that predates 5.19 may take the weak branch, and async_without_cancel_flags, which forces the no-flags path on every kernel, is unchanged (M1). Before this, a probe that got no answer on a 6.x runner — which N1 made a per-New event — silently turned that assertion into one any run with the reap off passes trivially, and the subtest is the registered deterministic killer of mutants m01 and m04.
  • The cancel-flags probe's four test seams carry the locking constraint they are read under (N-a); TestReapOnTheRunningKernel's two probes stay, because neither of its subtests weakens a check and probe_matches_the_kernel fails loudly on a disagreement (N-d); two comments round 4 left ragged are reflowed (N-b); and the residual below is round 4's new per-New cost (N-c).

Refs #657. PR 3 closes it.

What changes, and why each part exists

Everything is in engine/iouring:

  • a new file, fd_lifetime.go;
  • changes to transplant_source.go, worker.go, cqe.go, conn.go, handoff_loss.go, probe.go and engine.go.

Eight new EngineMetrics fields are added in engine/engine.go and summed over both sub-engines in adaptive/engine.go.

P1. The R0 gate at both hand-off sites. tryTransplant (sync) and finishAsyncTransplant (promoted async) refuse while recvArmed || kernelInflight != 0 || zcNotifPending. When the only op in flight is the recv, they start a REAP (P2). The dup, detach, close and adopt sequence the two sites shared is now one function, handOff, reached only after the gates.
Why: this is the precondition of every measured loss. On the base, TransplantHandoffInFlight equalled the number of hand-offs in every one-worker run.

P2. REAP, a reported cancel of exactly that recv.

  • prepCancelUserDataReported matches on the recv's user_data, generation included, so it cannot touch the next owner of the fd number. Its own tag is udTransplantReap = 0x09<<56.
  • The hand-off happens at the recv's own -ECANCELED, which handleRecv routes first (reapOutcome), before the generic res <= 0 close branch.
  • A request that arrives first is served here and its response is HELD (P3).
  • A cancel that reports it matched nothing (0 or -ENOENT) is retried on the next loop iteration while the recv is still armed. A miss is never followed by a hand-off.
  • Any other result is a cancel that failed, not an answer about the recv. It is counted in TransplantReapFailed, is not retried, and is never followed by a hand-off. (Round 2.)
  • The reap needs IORING_ASYNC_CANCEL_ALL, which exists from Linux 5.19; through 5.18 the kernel returns -EINVAL and leaves the recv armed.
    • A runtime probe, probeAsyncCancelFlags, submits exactly the reap's SQE against a user_data nothing carries. Round 3 (R2), round 4 (N2, N3): it answers one of four things. Accepted (res >= 0 or -ENOENT) turns the reap on. Rejected (-EINVAL) is the kernel's answer, and is what every kernel before 5.19 gives. Unexpected is a completion the probe does not recognise; its reason names the errno ("cqe.res=-9 (EBADF)"). No answer (the ring could not be set up, the submit or the wait failed, no completion, a foreign completion) says nothing about the kernel. The last three keep the reap off, and the zero value of the answer is "no answer", so an answer never set cannot turn it on. New logs a rejection at Info, an unexpected answer at Warn on every kernel, and no answer at Warn when the kernel's version is 5.19 or later (Info below), each with the reason and the kernel. The engine-selected line carries async_cancel_flags and async_cancel_probe. createWorkers copies the answer to every worker.
    • Round 4 (N1): only the kernel's answer, accepted or rejected, is cached for the process. After no answer, or an answer the probe does not recognise, the next New probes again and logs again, so one transient NewRing failure — EMFILE or ENOMEM while the adaptive engine builds its io_uring engine under load — no longer keeps the reap off for the life of the process; a mutex serialises the probes. The probe's wait for its cancel's completion is repeated once when it comes back early with nothing to read: SubmitAndWaitTimeout returns nil for a wait a signal cut short (EINTR) as well as for its timeout, so only an early return can be the former.
    • Where the reap is off, no reap is ever placed and no retry is queued; each refusal is counted in TransplantReapUnsupported. A connection the worker serves itself keeps its recv until that recv completes on its own, and its next response is then HELD, where the worker has no provided-buffer ring (HOLD needs none; a kernel that rejects the flags has none, since buffer rings also arrived in 5.19). With one it stays on io_uring.
    • Round 3 (R1): a promoted async connection is not offered for the hand-off at all on such a worker: asyncTransplantEligible refuses, so its dispatch goroutine parks as it does with no drain set, and it stays on io_uring (placement only). Before, the goroutine claimed at every park, the claim was refused, and the next request respawned it: a spawn and a detach-queue round trip per request for as long as the drain lasted.
  • A hand-off that fails at its dup (EMFILE) or at SetNonblock sets reapSuppressed: no reap is placed for that connection until it next receives data (round 2). The dup goes through a Worker.dupFD test seam. Round 3 (R5): only data and the release clear the flag, not the end of the drain, and the field's doc now says what that does across drains.
  • reapOutcome consumes the reaped recv's last completion, so it does what handleRecv does at any recv's last completion: a flagged provided buffer goes back to the ring and recvLinked is cleared (round 2, defensive).

Why a count and not a bool: this is the #484/#596 lesson. A stale miss must not clear the state a newer reap needs; mutant m05 shows it. Why its own tag: so a WebSocket pause's cancel (udRecvCancel, #596) is never taken for a reap, or the reverse.

P3. HOLD. While a drain is set and the worker has no provided-buffer ring, a connection the hand-off would accept has its response flushed unlinked, with no recv armed at all. Its SEND completion then finds nothing in flight and hands the connection off; the client's next request waits in the socket buffer for epoll. Both response tails go through one helper, respondAndArm. With no drain set, the SQEs are exactly those placed before: TestNoDrainSQESequenceIsUnchanged pins this, and it passes on the base too.

P4. The release guarantee.

  • releaseHold runs, unconditionally, after the transplant attempt at both SEND dispatch sites: the inlined one and processCQE's, which the listener-close harvest uses.
  • checkTimeouts rescues any held connection a path left unreleased and counts it in TransplantHoldRescued. This must stay 0.
  • The calls at the two RECV dispatch sites are removed (round 2, the m15 survivor): they were equivalent code, argued from the code and then measured. A held connection has no recv armed, because every recv-arming path clears the hold first or refuses a held connection. So the only recv completion that can reach such a call with the hold set is the one whose own response just set it, with its SEND still in flight, and releaseHoldSlow returns at its egress check. Measured on 86d46ac with a counting overlay, as tallied by the verifier's verify/tools/v4_tally.py over verify/v4/runs (verify/v4/V4-TALLY.txt; re-run in round 3 with identical output, its pr/V4-TALLY-regenerated.txt): the inline recv site was entered 60,559 times with the hold set and changed nothing, and processCQE's recv site 6 times, also nothing. The send site released 86 of its 90 entries in the same runs (the positive control).

P5. One owner per hand-off (A6, a pre-existing bug). tryTransplant leaves a promoted async connection alone when its dispatch goroutine has claimed its own hand-off (transplantPending, read under asyncInMu). finishAsyncTransplant acts only for the connState that still owns its slot. The two refusals are two counters (round 2):

  • TransplantClaimDeferred (a rate) is the first. It is ordering, not a double claim: a completion of the connection landed between the goroutine's park and the drain of its claim, and it is counted before any other gate.
  • TransplantDoubleClaim counts only a connState that no longer owns its slot. In async mode, only another hand-off of the same connection vacates the slot that way: a close marks the queued claim first, and a hijack is refused. So it must stay 0, and a release gate of 0 is meaningful.
  • A closing connection and the range check are plain refusals and are not counted.

P6. Submit pending SQEs before a worker parks (A5). When the worker is about to park indefinitely (no listener, no connections), it submits what the iteration queued. Pending close-path cancels were measured sitting unsubmitted at parks (1 to 24 of them) until something woke the worker. The comment says an idle SQ thread needs its NEED_WAKEUP kick; no tier enables SQPOLL today.

EngineMetrics fields (io_uring-only, cumulative, zero elsewhere; the adaptive engine sums both sub-engines):

field meaning expected
TransplantHeld responses flushed with the next recv held for a hand-off a rate
TransplantReaps reported cancels submitted for an armed recv a rate
TransplantReapMisses those that matched nothing (retried, never followed by a hand-off) a rate
TransplantReapFailed those that failed outright (not retried, never followed by a hand-off) must stay 0
TransplantReapUnsupported reaps not placed because the probe did not find the cancel flags accepted; only connections the worker serves itself are counted (a promoted async connection is never offered there) a rate, 0 wherever the probe finds the flags
TransplantHoldRescued held connections the timeout sweep had to re-arm must stay 0
TransplantDoubleClaim a hand-off refused because the connState no longer owned its slot must stay 0
TransplantClaimDeferred the sync path leaving a claimed async connection to its claim a rate

The gate this PR supports: W1 == 0 and W2 == 0, plus TransplantDoubleClaim, TransplantHoldRescued and TransplantReapFailed == 0. The rates are reported, not gated. Since round 3, T4 and the engine-level tests assert exactly this gate.

Tests

  • engine/iouring/fd_lifetime_test.go, synthetic completions through a ring-backed worker:
    • round 1: TestTransplantNeverHandsOffArmedRecv, TestTransplantReapMissIsRetried, TestHoldReleasedWhenDrainStops, TestHoldRescuedByCheckTimeouts, TestOneOwnerPerHandoff and TestNoDrainSQESequenceIsUnchanged;
    • round 2: TestTransplantReapFailureIsNotRetried, TestNoReapWithoutAsyncCancelFlags, TestReapSuppressedAfterFailedHandOff and TestReapedRecvLeavesNoLinkOrBuffer;
    • round 3, driving a real dispatch goroutine (asyncParkFixture, which records the goroutine each request ran on and sees a park in sync.Cond.Wait):
      • TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing (R1): three requests with the drain set on a worker without cancel flags; no claim, one goroutine for all three, nothing placed but the feed path's RECVs. Its control, control_async_park_claims_with_flags, shows the rig sees the claim, the reap and the hand-off where the worker can reap.
      • TestReapSuppressedAfterFailedHandOff/async_conn_retries_at_its_next_park: the async side of a failed dup (see R1 in the round-3 reply). reapSuppressed is never set at a park; the hand-off is retried once per request while the dup fails and the connection leaves once it works.
  • engine/iouring/async_cancel_probe_test.go and async_cancel_probe_log_test.go:
    • TestAsyncCancelProbeClassifies, with no_answer_is_not_a_rejection and log_levels (round 3), and round 4's zero_value_is_no_answer (N3) and unrecognised_answer_is_its_own_class (N2: the class, the errno in the reason, and the Warn on 5.15 and 6.8, for -EBADF, -ECANCELED and -EPERM);
    • TestAsyncCancelProbeOnThisKernel: logs the probe's answer beside the kernel version (round 3: no longer a verdict), drives the rejection path with a flag bit no kernel defines, and (round 4) a_cut_short_wait_is_retried_once, which holds the cancel back at the submit and cuts the wait short once, twice, and once at full length;
    • TestWorkersCarryTheAsyncCancelProbe, with no_answer_is_logged_apart_from_a_rejection (round 3: New with the probe's ring failing), and round 4's only_an_answer_is_cached (three calls per class of answer, counting the probes, then two News across a no-answer) and unexpected_answer_keeps_the_reap_off_and_is_not_cached;
    • TestReapOnTheRunningKernel, where the real kernel completes every op, and (round 3) probe_matches_the_kernel, which forces the flags on, lets the running kernel answer the real reap, and requires the probe's answer to be what the kernel did.
  • engine/iouring/fd_lifetime_engine_test.go, a real engine under keep-alive load: TestHandoffHasNothingInFlight, TestStaleRecvDataCounted (the join as a test: client errors == W1), TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen and TestWorkerParksWithNothingPending. Round 3: every load run fails on a non-zero TransplantDoubleClaim, TransplantReapFailed or TransplantHoldRescued. Round 4 (N4): where the engine's probe did not find the cancel flags, TestHandoffHasNothingInFlight/async wants none of the promoted connections handed off — such a connection is never offered for the hand-off, so it stays on io_uring — with the loss checks unchanged; it used to want all of them on every kernel, and so could not pass on 5.15.0-191. A new case, async_without_cancel_flags, takes the probe's answer as a rejection, so that branch runs on every kernel, the runner's included.
  • adaptive/flap_conns_per_ring_test.go: TestFlapConnsPerRing is T4. It uses two workers per engine, so its 256 keep-alive connections are 128 per io_uring ring, and runs three promote/revert cycles under load. It requires 0 client errors, and since round 3 (R4), from the adaptive engine's Metrics(): W1 = 0, W2 = 0, TransplantDoubleClaim, TransplantHoldRescued and TransplantReapFailed = 0, and at every switch that the engine switched away from detached at least all 256 and holds none while the one switched to adopted at least all 256 and holds all of them. Its memlock, io_uring-availability and adaptive.New skips all fail instead under CELERIS_REQUIRE_UPSWITCH=1.

Failing first, round 3. Source: round 3's ff/VERDICT.txt, from its tools/ff_verdict_r3.py over tools/run_ff_r3.sh's logs.

  • TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing (R1), on 48dfe1a plus only this head's two fd-lifetime test files: "the dispatch goroutine claimed the hand-off at 3 of 3 parks, and the 3 requests ran on 3 goroutines, want 0 claims and 1 goroutine", and "TransplantReapUnsupported = 3, want 0". It is the only failing subtest there; its control and the async suppression pin pass on 48dfe1a too.
  • TestWorkersCarryTheAsyncCancelProbe/no_answer_is_logged_apart_from_a_rejection (R2), on 48dfe1a plus only the probe-ring seam (round 3's ff/ovl/FF-48-R2SEAM/probe.go.seam.diff): "a probe that got no answer on kernel 7.0 was logged at INFO as "async cancel flags runtime probe failed: …", want WARN". The other R2 subtests use the new result type, which 48dfe1a does not have.
  • T4 (R4), alone at 128 MiB with io_uring refused by seccomp and CELERIS_REQUIRE_UPSWITCH=1: 48dfe1a's SKIPs ("io_uring unavailable: needs both sub-engines"), and this head's FAILs ("… -- CELERIS_REQUIRE_UPSWITCH=1 forbids skipping").
  • On this head all six tests and their 11 subtests PASS.
  • R3's and R5's changes are tests and docs. R4's new assertions are shown to fire by mutants r12 to r20 (below).
  • Round 2's failing-first results on 86d46ac are unchanged (round 2's ff/VERDICT.txt).

Failing first, round 4. Source: ff/VERDICT.txt, from tools/ff_verdict_r4.py over tools/run_ff_r4.sh's logs. Each of round 4's four checks is run on 272bcba plus only round 4's test seams (runAsyncCancelProbe, and the probe's submit, wait and wait timeout: ff/ovl/FF-272-SEAM/probe.go.seam.diff), under the subtest name this head gives it:

  • zero_value_is_no_answer (N3): "the zero value of asyncCancelProbe reads as accepted: an answer that was never set would turn the hand-off's reap on".
  • unrecognised_answer_is_its_own_class (N2): "classifyAsyncCancelProbe(-9) = rejected (…), want a class of its own, "unexpected"", and the record logged at INFO where WARN is wanted.
  • only_an_answer_is_cached (N1): "three calls with the probe answering no answer ran it 1 time(s), want 3", and "the first New's probe got no answer, the second's is accepted: asyncCancelFlags false then false, probes run 1; want false, true and 2".
  • a_cut_short_wait_is_retried_once (N1): "wait cut short once: the probe gave (no answer, "no CQE produced for the cancel (waited 500ms)") after 1 wait(s), want accepted after 2".
  • The seam patch alone leaves 272bcba's own four probe tests passing (arm FF-272-SEAMCTL), so the seams change no behaviour, and the full-length-wait check — which is not a behaviour change — does not fire there.
  • On this head the four probe tests and their 11 subtests PASS.

CI

ci.yml is unchanged since round 2. The new tests are subtests of the 29 tests the witness step names, so its tally and its ban on any SKIP line (indented ones too) cover them.

  • unit job, witness step. It names 29 tests: PR 1's 11 and PR 2's 18. It runs with -v, CELERIS_REQUIRE_IOURING_WORKERS=1 and an exact tally of top-level === RUN and --- PASS: Name ( lines, and it fails on any --- SKIP line. Every celeris657 engine workers=N line must read 1.
  • adaptive job, race step (memlock raised). It tees its log and tallies T4: exactly one top-level RUN and one PASS, no SKIP line for T4, and exactly one S0T4 RESULT line reading conns=256 wE=2 wI=2 conns_per_ring=128 … err=0. The RESULT line now also carries w1= w2= doubleclaim= holdrescued= reapfailed= held= reaps=, after err=, so the tally's pattern is unchanged.
  • adaptive job, one-worker leg (runner memlock). Exactly 2 RUN and 2 PASS, no SKIP line, every io_uring engine listening line at workers=1, and each test's own summary line at err=0. The verdicts tolerate up to 64 lost requests; a base flap run PASSED with 56.

Departures and what they imply.

  • The design asked for Resources.Workers: 1 variants of the revert tests. resource.MinWorkers is 2 (resource/config.go:147-148), so they cannot be built.
  • The one-worker shape instead comes from the runner's 8 MiB memlock: io_uring needs 12 MiB per worker (minMemlockPerWorker), so it starts one.
  • So if GitHub raises the runner's default memlock, the witness step and the one-worker leg turn red, because they require workers=1. That is intended: the shape must then be pinned another way.
  • T4 does not depend on this. It fixes two workers explicitly and runs at raised memlock.

Every tally, proven both ways. Round 2 ran 28 arms at 48dfe1a from the verbatim run: blocks (bash --noprofile --norc -eo pipefail, a sudo shim, one golang:1.27 container per arm) and all 28 proved their side (round 2's interlock/48dfe1a/VERDICT.txt, tools/il_verdict_r2.py). With ci.yml unchanged, round 3 re-ran the arms whose run changed. Source: round 3's interlock/1f883e2/VERDICT.txt, from its tools/il_verdict_r3.py:

arm step result
committed tree witness step pass: 29/29, SKIP 0, engines 7 at workers=1: 7
committed tree one-worker leg pass: 2/2, SKIP 0, engines 2 at workers=1: 2, err 0 and 0
committed tree race step pass: T4 1/1, RESULT wE=2 wI=2 … err=0; package 76 PASS, 0 FAIL, 0 SKIP
verdicts neutralised, fix engine race step pass: RESULT err=0
io_uring refused by seccomp race step fail: T4 now FAILs ("forbids skipping"), so go test fails. In round 2 T4 skipped, go test was ok, and the tally refused the SKIP.
T4 edited to 3 workers race step fail: T4 PASSes, and the tally refuses RESULT wE=3 wI=3
no-HOLD/REAP engine race step fail: T4 FAILs, err=167, with its LOSS, W1, W2 and MOVE messages
verdicts neutralised, no-HOLD/REAP engine race step fail on the RESULT line alone: go test ok, err=158

8 of 8 arms prove their side, each in its own container, one at a time.

Round 4 re-ran the three GOOD arms on this head, from the same verbatim run: blocks, one container each (interlock/5ce4986/VERDICT.txt, tools/il_verdict_r4.py). ci.yml is byte-identical to 272bcba's (git diff 272bcba HEAD -- .github/ is empty), so the other arms are round 3's and stand.

  • witness step: pass, 29/29, SKIP 0, engines 8 at workers=1: 8 (one more engine than round 3: the new async_without_cancel_flags case builds one).
  • one-worker leg: pass, 2/2, SKIP 0, engines 2 at workers=1: 2, summaries err 0 and 0.
  • race step: pass, T4 1/1, RESULT cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=2109529 err=0, package 0 FAIL, 0 SKIP.

On the runner, both ways (R6). Run 35465814564 was red on purpose. Its commit, 683c7bb ("DO NOT MERGE"), planted the no-HOLD/REAP engine and turned T4's and the revert tests' verdicts into logs, so only the steps' own tallies could fail. From its real job logs (round 3's ci/683c7bb/SUMMARY.txt):

  • The race step's go test passed (76 PASS). Its T4 tally refused the RESULT line (err=188 w1=188 w2=398, "at two workers with err=0: 0").
  • The one-worker leg's two tests passed. Its tally refused their summaries (err 13 and 36, "summaries 2 with err=0: 0").
  • The unit job (the engine-level tests on that engine) and lint (the functions the plant left unused) went red too.

Commit 272bcba reverts it; its tree is identical to 1f883e2's. Run 35465988505 on it is green. Both runs are in round 3's RED-RUNS.txt.

How it was checked

Mutation (round 4). Tool: tools/mkmut_r4.py; runner: tools/run_mut_r4.sh; verdict: mut/VERDICT.txt, from tools/analyze_mut_r4.py; the re-scoped mutants are explained in mut/VERDICT-NOTE.txt and mut/EXPOSURE.txt (tools/exposure_census_r4.py).

  • 85 mutants, one container each: round 3's 74 re-derived on this head, and 11 new ones, q01 to q11. They run in three profiles:
    • the 29-test witness list at 8 MiB with CELERIS_REQUIRE_IOURING_WORKERS=1 (75 mutants);
    • the adaptive metrics tests at 8 MiB (4);
    • T4 alone at 128 MiB with CELERIS_REQUIRE_UPSWITCH=1 (6).
  • Round 3's new mutants (r01 to r20) covered R1 to R4: claims without the flags and the check inverted; the "sticky" suppression a park-side check would need; no answer read as a rejection, logged at Info, or taken as accepted by New; the 5.19 threshold off by one; the ring seam ignored; a probe that always says rejected; and each counter that must stay 0 planted against the engine-level tests and against T4, plus the reverse drain turned off.
  • Round 4's (q01 to q11): the cache keeping every answer, keeping the unrecognised one, or keeping none; no wait retry, a retry after a full-length wait, and a second retry; an unrecognised answer folded back into a rejection, logged at Info, turning the reap on, or with the errno dropped from its reason; and the old constant order, where the zero value was "accepted".
  • All 85 are killed, each by the tests named for it, and for the registered reason where one is. The unmutated controls passed 29/29 with 44 subtests, 2/2 and 1/1. Every io_uring run printed workers=1 for all 8 engines, and every T4 run's RESULT line read wE=2 wI=2. 0 survivors.
  • Three round-1 mutants are re-scoped (N5), and only their named killers changed. The review's point about m04 is that TestStaleRecvDataCounted reaches the branch m04 changes only if a connection's recv happens to be armed when the drain examines it. Across the runs of rounds 2 to 4 that exposure ranges from 0 to 88 of 128 and is 0 in 6 of 18 (mut/EXPOSURE.txt, tools/exposure_census_r4.py), so its FAIL was never guaranteed.
    • m04 and m01 (the same gate, removed rather than bypassed) now name the two deterministic killers: TestTransplantNeverHandsOffArmedRecv, which drives the branch with synthetic completions, and TestHandoffHasNothingInFlight/async, where every connection is promoted and a promoted connection always leaves through the reap (W2 = 128 in both m04 runs).
    • m23 (the same gate removed with the witnesses blinded) can be seen by the engine-level tests only through their clients' errors, which needs the same exposure. It now names the three tests whose checks read the hand-off itself: TestTransplantNeverHandsOffArmedRecv and TestTransplantHandoffInFlightCounts{TryTransplant,FinishAsyncTransplant} ("the R0 gate is gone").
    • TestStaleRecvDataCounted still runs in every one of those mutants' -run sets, and mut/VERDICT.txt reports what it did; it is simply not required.

Whole-package suites, this head against 272bcba. One container per tree per cell, with -race -v, count=1. Source: suites/COMPARE.txt, from tools/compare_suites_r4.py over tools/run_suites_r4.sh's logs.

cell this head, top-level 272bcba, top-level
./engine/iouring, 8 MiB, as the unit job runs it 154 PASS, 5 SKIP 154 PASS, 5 SKIP
./engine/iouring, 128 MiB, CELERIS_REQUIRE_IOURING_WORKERS=1 159 PASS 159 PASS
./adaptive with the CI -skip, 8 MiB 72 PASS, 4 SKIP 72 PASS, 4 SKIP
./adaptive, 128 MiB, CELERIS_REQUIRE_UPSWITCH=1 76 PASS 76 PASS
  • 0 items PASS on 272bcba and fail to PASS on this head. None of the 8 runs has a FAIL, a WARNING: DATA RACE or a panic, and every 8 MiB run's io_uring engines are at one worker.
  • The only differences are round 4's six new subtests in ./engine/iouring, absent on 272bcba and PASS on this head in both memlock cells: zero_value_is_no_answer, unrecognised_answer_is_its_own_class, only_an_answer_is_cached, unexpected_answer_keeps_the_reap_off_and_is_not_cached, a_cut_short_wait_is_retried_once and TestHandoffHasNothingInFlight/async_without_cancel_flags.
  • tools/t4_switch_census_r4.py → t4-switch-census.txt covers this round's unmutated T4 runs (the interlock T-GOOD arm, the mutation up128 control and the 128 MiB adaptive suite cell). All 18 switches moved and adopted 256 connections, and all 3 RESULT lines read err=0 w1=0 w2=0 doubleclaim=0 holdrescued=0 reapfailed=0.

Lint, as the lint job runs it. golangci-lint v2.13, with the repo's config, is clean in all five modules on linux/amd64 and linux/arm64. mage -compile and mage CheckRelease pass. actionlint 1.7.12 with shellcheck is clean, and so is zizmor 1.30.0 (offline). Each linter flags a defect planted in a copy of the tree (lint/, tools/run_lint.sh; re-run on this head). gofmt lists one file, test/benchcmp_ws/bench_test.go, which has been untouched since #262 and is listed on 272bcba too; no CI step runs gofmt.

GitHub CI on this head. Runs 35474330959 (CI) and 35474327611 (this PR's analysis) on 5ce4986: ubuntu-24.04, linux/amd64, kernel 6.17.0-1022-azure, go1.27.0. All 11 jobs were green at the first attempt, and none was re-run. From the real job logs (ci/5ce4986/SUMMARY.txt, tools/summarize_ci_r4.py, which counts worker witnesses only from engine and test log lines):

  • Witness step, memlock 8192 KiB: want 29, ran 29, passed 29, SKIP lines 0, engines 8 at workers=1: 8, with 44 subtests PASS (round 3: 38; round 4's six are the difference). The probe answered accepted on the runner's kernel and the forced reap ran there (probe_matches_the_kernel), and async_without_cancel_flags PASSed, having driven the no-flags branch on that kernel: "0 of 128 promoted async conns handed off (promotions 128)".
  • Adaptive race step, memlock unlimited: 76 PASS, 0 FAIL, 0 SKIP, and celeris#657 T4: want 1, ran 1, passed 1, SKIP lines 0, RESULT lines 1, at two workers with err=0: 1. The RESULT line reads cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=616472 err=0 w1=0 w2=0 doubleclaim=0 holdrescued=0 reapfailed=0 held=750 reaps=140, and every one of the six switches moved and adopted 256 connections.
  • One-worker leg: want 2, ran 2, passed 2, SKIP lines 0, io_uring engines 2 at workers=1: 2, summaries 2 with err=0: 2.
  • Other jobs: the io_uring: a worker whose ring setup or first submit fails leaves its SO_REUSEPORT listen socket open, and adaptive retries the build every tick (found by reading) #656 init-failure tests passed 3/3 and conformance 4/4 (115 subtests).

The witness tests on real Ubuntu kernels under QEMU, at this head (N4). Round 2's rig, the same Ubuntu-archive .debs — tools/qemu_inner_r4.sh compares their hashes with round 2's verifier before booting and refuses if they differ (they are identical) — qemu-system-aarch64 (TCG), and the head's ./engine/iouring test binary built with -race. Source: qemu/VERDICT.txt, from tools/qemu_verdict_r4.py; tools/qemu_plans_r4.sh, qemu_build_r4.sh, qemu_run_r4.sh + qemu_inner_r4.sh.

kernel the probe's answer the kernel executed the forced reap TestReapOnTheRunningKernel/probe_matches_the_kernel
5.15.0-191 rejected: cqe.res=-22 (EINVAL) no PASS, on its rejecting branch: the kernel failed the reap (TransplantReapFailed=1) and the connection left after its held response (TransplantHeld=1, adopted 1)
5.19.0-50 accepted yes PASS
6.8.0-138 accepted yes PASS
  • On all three kernels the five probe tests PASS, round 4's own subtests PASS, every engine runs at one io_uring worker, and there is no WARNING: DATA RACE. This is the first time the rejecting branch has run anywhere.
  • The run also found TestHandoffHasNothingInFlight/async failing on 5.15 — "0 of 128 busy conns were handed off, want all" — on 272bcba and on the first round-4 head alike, because a promoted async connection is never offered for the hand-off where the flags are missing. 5ce4986 fixes the expectation; the subtest passes on 5.15 now, and async_without_cancel_flags runs that branch on every kernel.
  • Two failures under QEMU are the apparatus, not this head: 272bcba's binary has them too. The follow-up that shows it is qemu/diag/ (tools/qemu_build_diag_r4.sh, qemu_run_diag_r4.sh), which runs the same two halves for 272bcba's -race binary, the head without -race, and the head with a diagnostic 10 s client read deadline (never committed).
    • TestReapedRecvLeavesNoLinkOrBuffer/provided_buffer_goes_back on 5.15 and 6.8: "register pbuf_ring: invalid argument", no provided-buffer ring in those boots. 272bcba fails the same subtest on the same two kernels.
    • The engine-level load tests' read_timeout errors on 5.15 and 5.19: TCG emulates the machine in software. Without -race all four pass on 5.15, and with the 10 s client read deadline they pass on 5.15 and 5.19; 272bcba fails them there too. On 6.8 all four pass as they are.
  • Round 2's QEMU run stands but is superseded by this one, which is on this head and with -race.

The campaign (round 1, on 308ffca), unchanged. n = 8 runs per arm. W1 is StaleRecvDataTransplanted + Unattributed, and p is Fisher's exact test, two-sided, on runs that FAIL. Every figure comes from ship/tools/regate.py over the raw campaign logs (ship/regate/REGATE.txt), which imports nothing from the campaign's own analyzer.

cell base 5b2e83b fix p
TestReverseTransplant, 1 io_uring worker, 64 conns FAIL 4/8; 56 lost, == W1 in 8/8; W2 = 512 of 512 hand-offs FAIL 0/8; 0 lost; W1 = W2 = 0; 512 hand-offs (512 held, 1 reap, 1 miss) 0.077
TestBidirectionalFlap, 1 io_uring worker, 64 conns FAIL 7/8; 386 lost, == W1 in 8/8; W2 = 512 of 512 FAIL 0/8; 0 lost; W1 = W2 = 0; 512 hand-offs (505 held, 8 reaps, 1 miss) 0.0014
T4 (TestFlapConnsPerRing: 2 io_uring workers, 256 conns, 3 cycles) FAIL 8/8; 1,713 lost, == W1 in 8/8; W2 = 3,030 of 3,030 FAIL 0/8; 0 lost; W1 = W2 = 0; 6,144 hand-offs (6,123 held, 21 reaps, 0 misses) 0.00016
switch-stall harness (TestSwitchDiag662: 1 io_uring worker, 4 switches per run), n = 60 stall rounds 18/60; 63 keep-alive failures; 13 churn read timeouts; W1 = 76; W2 = 3,867 of 3,867 stall rounds 0/60; 0 keep-alive failures; 0 churn read timeouts; W1 = W2 = 0; 3,911 hand-offs (3,765 held, 151 reaps, 0 misses) 1.7e-06
  • On the switch-stall harness, W1 equals keep-alive failures plus churn read timeouts in 60 of 60 base runs (76 = 63 + 13). This is computed by the verifier's verify/tools/v7_e5_w1_join.py; no campaign tool computed it.
  • The same harness's accept-churn clients saw these non-OK results over the 60 runs (verify/tools/v7_e4_churn_census.py):
    • fix: reset 189, refused 16, dial reset 3;
    • base: reset 179, refused 1, dial reset 4, timeout 13.
  • Negative controls, where the fix has one part removed and must lose:
    • without HOLD/REAP: TestReverseTransplant FAIL 5/8 and TestBidirectionalFlap FAIL 7/8, with lost == W1 in every run;
    • without HOLD/REAP and with W1 not counted: FAIL 7/8 and 7/8, with W1 and W2 reading 0, so the tests still catch the loss by its client errors alone;
    • without A6: TestOneOwnerPerHandoff FAIL 6/6;
    • without A5: TestWorkerParksWithNothingPending FAIL 6/6.
  • Phase B compared nine whole-package cells, n = 2 per arm, and found no item FAILing more on the fix than the base beyond its cell's A/A floor.
  • The campaign has 438 container logs plus a build log, and none contains WARNING: DATA RACE or a panic: line (verify/tools/v7_e5_logcount.sh).
  • The worker-count rule, corrected in round 2 (E3): tools/validity_r2.py reads both worker witnesses, and tools/e3_rerun.py re-ran both rules over all 308 phase B, C and S logs (tooling/E3-RERUN.txt): 0 verdicts change, the 26 unit arms that build an engine all read workers=1, and regate's worker check finds 0 violations.

Interactions with other open issues

Tracked elsewhere

What this PR does not do

This is PR 2 of 3:

Known residuals

  1. Where the probe does not find the cancel flags (every kernel before 5.19): the hand-off never reaps there. A connection the worker serves itself leaves after its next, held, response. A promoted async connection is not offered for the hand-off and stays on io_uring after a revert, parked (round 3; before, it stayed too, but its goroutine was respawned per request). That is a placement cost, never a loss, and PR 3's sweeps own placement. Adaptive reaches io_uring below 6.1 only with CELERIS_ADAPTIVE_START=iouring. The rest of that kernel range's cancel breakage is io_uring: every async cancel fails with -EINVAL on Linux 5.10-5.18 (IORING_ASYNC_CANCEL flags are 5.19+) #682.
  2. A stale reap count can outlive an -EALREADY'd cancel. The effect is placement only: a later miss queues no retry, so the connection stays until its client sends again, and then HOLD hands it off.
  3. Reap misses are frequent under continuous load in the engine-level tests, up to one miss per reap (for example 114 of 114 on 6.8.0-138 in round 2's QEMU run). Every one resolved safely, with 0 lost and W2 = 0.
  4. Multishot recv (CELERIS_IOURING_MULTISHOT_RECV=1, opt-in) disables HOLD, so only REAP applies there. No test exercises it.
  5. Refused dials at a revert: measured, pre-existing, now adaptive: a revert can refuse dials — epoll ResumeAccept returns before any loop listens, then io_uring closes its listener (3.4% of reverts) #683. A pre-registered measurement on round 2's head (measure/00-PREREG.md, 104 runs per arm, the verdict computed by measure/tools/analyze.py into results.json; the comment fix(iouring): hand a connection to epoll only when no read can still resolve its fd (celeris#657) #681 (comment)) found no regression:
  6. One host shape: all local evidence is linux/arm64 on one Docker VM, plus round 2's QEMU kernels. GitHub CI adds linux/amd64 runs. Both cluster architectures (nightly and 24 h soak, with the validator gating on the counters above) are owed before release.
  7. A cancel-flags probe that gets no answer is re-run at every New. Round 4 (N1) caches only the kernel's answer, accepted or rejected, so a kernel that never answers — or answers with something the probe does not recognise — is probed again at every iouring.New, serialised on a process-global mutex, with a log record each time (WARN from 5.19, INFO below it). Each probe opens and closes a private ring; in the remote case where the cancel produces no CQE at all, it holds that mutex for up to 500 ms (a wait to a 500 ms deadline, repeated once if a signal cuts it short). The price buys back the failure mode it replaced: cached, one EMFILE or ENOMEM in that private ring kept the hand-off's reap off for the life of the process. No kernel in this PR's evidence takes the path — on every one of them the first probe answers, is cached, and the cost is one probe per process — and from round 5 a 5.19-or-later kernel where it happens fails TestHandoffHasNothingInFlight/async rather than quietly weakening it (M1).

…etime (celeris#657)

PR 2 of 3 for #657 (face 2). These tests fail on main (5b2e83b) and pin
what the fix must do; the fix follows in separate commits.

Unit tests drive one connection through the real request, send and
hand-off code with the kernel taken out of the loop: every SQE the engine
places is read back and removed unsubmitted, and every completion is
written by the test, so each ordering a cancel can resolve in is a
deterministic input.

- TestTransplantNeverHandsOffArmedRecv: no hand-off while the recv is
  armed; a reported cancel of its own tag (0x09) instead, hand-off at the
  recv's -ECANCELED, a request that beats the cancel served and HELD, and
  the same against the real kernel.
- TestTransplantReapMissIsRetried: a missed reap is retried, never
  followed by a hand-off; a stale miss cannot clear a newer reap.
- TestHoldReleasedWhenDrainStops: a held conn gets its recv armed at the
  SEND completion when it is not handed off (processCQE, the worker loop,
  a refused hand-off).
- TestHoldRescuedByCheckTimeouts: the timeout sweep's belt, counted.
- TestOneOwnerPerHandoff: a conn claimed by its dispatch goroutine is not
  also moved by tryTransplant; finishAsyncTransplant leaves a slot it no
  longer owns alone.
- TestNoDrainSQESequenceIsUnchanged: with no drain set the per-request
  SQEs (opcode, flags incl. IO_LINK, tag) of the sync, WRITEV and
  direct-body tails. Passes on main; it is the witness that the fix
  changes nothing on the steady-state path.

Engine-level (real worker loop, keep-alive clients across a
StartTransplant): TestHandoffHasNothingInFlight, TestStaleRecvDataCounted
(client losses == stale data CQEs, and both 0),
TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen and
TestWorkerParksWithNothingPending.

Refs #657
…sites (celeris#657)

tryTransplant and finishAsyncTransplant each carried their own copy of the
commit sequence (dup, detach, witnesses, cancel, release, close, adopt).
Both now call handOff once their gates have passed, so the fd-lifetime
gate that follows is written once and the celeris#657 witnesses stay
where a gate regression would move them. No behaviour change: the order
of every step and the release kind per site are as before.

Refs #657
…celeris#657)

EngineMetrics gains five io_uring-only, cumulative fields, exported the
way the PR-1 witnesses are (the engine's handoffLoss set, Metrics(), and
the adaptive sum over both sub-engines):

- TransplantHeld: responses flushed with the next recv held for a
  hand-off (HOLD).
- TransplantReaps / TransplantReapMisses: reported cancels of an armed
  recv submitted so the conn can be handed off (REAP), and those that
  matched nothing.
- TransplantHoldRescued: held conns the timeout sweep had to rescue.
  Must stay 0.
- TransplantDoubleClaim: hand-offs refused because another path owned
  the conn's hand-off (a double hand-off prevented).

Nothing increments them yet; the rule that does follows.

Refs #657
…ve its fd (celeris#657)

P1 (R0 gate) and P2 (REAP) of the fd-lifetime rule. Both hand-off sites,
tryTransplant and finishAsyncTransplant, now refuse a connection with
recvArmed, kernelInflight != 0 or zcNotifPending. When the only op in
flight is the recv, they reap it: a REPORTED cancel matched on the recv's
own user_data (generation included) under a new tag, udTransplantReap =
0x09<<56. The hand-off runs again at the recv's -ECANCELED, which
handleRecv now routes first, before the generic negative-result branch
that closes the conn. finishAsyncTransplant reaps instead of cancelling,
dup'ing and closing with the recv armed.

- If the request arrives first, the recv completes with it and it is
  served here as usual.
- A cancel that reports a miss is retried on the next loop iteration
  while the recv is still armed (drainDetachQueue runs the retries). A
  miss is never followed by a hand-off.
- The per-connection state is a COUNT of outstanding reaps plus a flag
  saying they all target recvs that have since completed, so a newer recv
  can get its own reap and a stale miss cannot clear it (the
  celeris#484/#596 lesson).
- A pause's cancel keeps precedence in the -ECANCELED routing; detached
  WS/SSE conns are never reaped.

The PR-1 witness tests drove the sites with a recv armed, which the gate
now refuses. They check the refusal and then drive handOff, the commit
point the gate guards, so the witnesses stay proven.

Refs #657
…s set (celeris#657)

P3 (HOLD). While a drain is set and there is no provided-buffer ring, the
response of a connection the hand-off would accept is flushed unlinked and
its next recv is not armed at all. That SEND's completion then finds
nothing in flight and hands the connection off; the client's next request
waits in the socket buffer for epoll.

Every recv-arming tail after a response now goes through one helper,
respondAndArm: the sync tail (flushSendLink, or flushSend on a buffer
ring) and the direct-body tail (flushSend, then a standalone recv). With no
drain set it places exactly the SQEs the two tails placed before
(TestNoDrainSQESequenceIsUnchanged); the added cost per request is one
atomic pointer load.

Held connections the hand-off does not take are released by the next
commit.

Refs #657
…e (celeris#657)

P4, the release guarantee for HOLD. A held connection has no recv armed,
so if the hand-off at its SEND completion does not happen (the drain
stopped, a gate refused, the dup failed) nothing would ever read its next
request.

- releaseHold runs after the transplant attempt at both inlined dispatch
  sites, unconditionally: once nothing is left to send, the hold is
  cleared and the recv armed.
- processCQE's udRecv and udSend cases gain the transplant attempt they
  lacked, and the release. The listener-close harvest handles completions
  there.
- checkTimeouts rescues a held connection found with its send done and
  arms its recv, counted in TransplantHoldRescued, which must stay 0.

Refs #657
P5 (A6). A promoted async conn whose dispatch goroutine had parked and
claimed its own hand-off (transplantPending, asyncRun cleared) could be
moved a second time: a SEND completion landing before the worker drained
the claim reached tryTransplant, which only skipped a RUNNING goroutine
and moved the conn as if it were sync; drainDetachQueue's
finishAsyncTransplant then dup'd cs.fd, a number the first move had closed
and the process had reused. Measured on main: one identity moved twice,
20 us apart, in 1 of 167 runs, joined to the only negative gauge.

- tryTransplant reads the claim with asyncRun under asyncInMu (the lock
  the goroutine sets both under) and refuses a claimed conn.
- finishAsyncTransplant returns unless w.conns[cs.fd] == cs and the conn
  is not closing.

Both refusals count in TransplantDoubleClaim.

Refs #657
P6 (A5). The iteration that closes or hands off a worker's last
connection queues that connection's close-path cancels in the same pass
that finds the worker idle, and the DRAINING->SUSPENDED park is
indefinite, so the worker parked with them unsubmitted (measured 1-24
pending SQEs at parks). It now submits before taking wakeMu and setting
suspended.

Refs #657
…ting target (celeris#657)

The warm-up response's SEND completion can reach the worker after the
client has read the response, and so after StartTransplant. The idle conn
is then reaped and handed to the target before /big is sent; the test's
accepting target closed it, and the client read a reset (8 of 15 runs on
17ab243, reported with adopted=1 held=0). The target now refuses, so such
a conn is reclaimed onto io_uring and /big is served and held there as the
test intends; it also asserts the response really was held, and reports
the hand-off counters when the read fails.

Refs #657
…en, on one worker (celeris#657)

The unit job runs ./engine/iouring without -v, so a skip there is
invisible. The ten fd-lifetime tests run again by name with -v and
CELERIS_REQUIRE_IOURING_WORKERS=1, with the witness step's exact tally
(RUN and PASS counts, no SKIP line). It is also the one-worker leg: the
runner's 8 MiB memlock funds a single io_uring worker, the shape the loss
was measured in, and the step fails unless every engine the tests start
logs workers=1.

Refs #657
…fd-lifetime counters (celeris#657)

TestHandoffHasNothingInFlight gains an async subtest: every route async,
so each conn is promoted to a dispatch goroutine, claims its own hand-off
at its park and leaves through finishAsyncTransplant, which must reap the
recv the feed path armed. The load helper logs the new counters (held,
reaps, misses, rescued, double claims) next to the witnesses.

Refs #657
…n the revert tests on one worker (celeris#657)

The unit job's celeris#657 witness step now carries all twenty-one tests:
the eleven hand-off loss witnesses from PR-1 and the ten fd-lifetime tests
this PR adds. It runs them by name with -v, CELERIS_REQUIRE_IOURING_WORKERS=1
and an exact tally (twenty-one RUN, twenty-one PASS, no SKIP line). Each
engine-level test logs its io_uring worker count, and the step requires
every such line to read workers=1 at the runner's own 8 MiB memlock, the
one-ring shape in which the hand-off losses were measured.

The adaptive job gains a second step for TestReverseTransplant and
TestBidirectionalFlap at the runner's memlock. The job's existing step
raises memlock, so until now those two tests never ran with io_uring on a
single worker. The new step forbids skipping (neither test reads a REQUIRE
variable, so the tally fails on any SKIP line), requires exactly two RUN
and two PASS, and requires every "io_uring engine listening" line to read
workers=1. It runs even when the step before it failed.
@FumingPower3925 FumingPower3925 added this to the v1.6.0 milestone Sep 19, 2026
@FumingPower3925 FumingPower3925 added bug Something isn't working engine/iouring io_uring engine specifics labels Sep 19, 2026
…plit the double-claim counter (celeris#657)

Round 2 of #681 (PR 2 of 3 for #657). Review findings C1-C4, C6 and the
m15 survivor.

C1. The reap cancels an armed recv with IORING_ASYNC_CANCEL_ALL, and
cancel flags exist only from Linux 5.19. Through 5.18 the kernel fails
the cancel with -EINVAL and leaves the recv armed (measured on
5.15.0-191), and handleTransplantReap read -EINVAL as a miss and
requeued it: one failing cancel per loop iteration, per connection, for
as long as the drain lasted.
- New probeAsyncCancelFlags (cached, like the other runtime probes)
  submits exactly the reap's SQE form against a user_data nothing
  carries: res >= 0 or -ENOENT means accepted, -EINVAL rejected. New
  stores the answer, logs it (async_cancel_flags), and createWorkers
  copies it to every worker.
- startReap places no reap where the flags are rejected and queues no
  retry. The connection keeps its recv until the recv completes on its
  own; a sync connection is then HELD at its next response and leaves at
  that SEND's completion with nothing in flight. Counted in the new
  TransplantReapUnsupported (a rate, 0 from 5.19).
- handleTransplantReap retries only a miss (0, -ENOENT). Any other
  non-hit result is counted in the new TransplantReapFailed (must stay
  0) and is neither retried nor followed by a hand-off.
The other cancel sites have the same pre-5.19 problem on main; they are
not changed here.

C2. A hand-off that fails at its dup (EMFILE) left the conn in place,
and the recv re-armed for it was reaped again at once: a RECV, a cancel
and two completions per iteration while the failure lasted. handOff now
sets reapSuppressed on a dup or SetNonblock failure; startReap places
nothing while it is set; handleRecv clears it when data arrives. The dup
goes through a Worker.dupFD test seam (nil = unix.Dup).

C3. TransplantDoubleClaim counted two different things. tryTransplant
finding a conn whose dispatch goroutine had claimed its own hand-off is
ordering (a completion landed between the park and the drain of the
claim); it is now TransplantClaimDeferred, a rate. finishAsyncTransplant
counts TransplantDoubleClaim only when the connState no longer owns its
slot, which in async mode means another hand-off of the same conn; a
closing conn and the range check are plain refusals. So a release gate
of TransplantDoubleClaim == 0 now means what it says.

C4. reapOutcome consumes the reaped recv's last completion, so it now
does what handleRecv does at any recv's last completion: it returns a
flagged provided buffer to the ring and clears recvLinked.

C6. The A5 comment said SQPOLL "submits by itself"; an idle SQ thread
needs the NEED_WAKEUP kick. Comment corrected; no tier enables SQPOLL.

m15 (removing releaseHold after the inlined udRecv dispatch) survived
round 1's mutation run. It is an equivalent mutant, so the call is
removed, together with the same call in processCQE's udRecv case, which
has the same status (round 1's m12 was killed only through the send
site). Proof, from code at 86d46ac:
- transplantHold is set in one place, respondAndArm, only with no
  provided-buffer ring (single-shot: the recv that carried the request
  has just completed, nothing is armed), and only when holdEligible's
  last clause holds (cs.sending or bytes pending). flushSend never
  completes a send synchronously.
- Every recv-arming path clears the hold first (releaseHoldSlow,
  rescueHold) or refuses a held conn (reapOutcome), and detached,
  paused, H2 and promoted async conns are never held.
- So a held conn has no recv armed, and the only recv completion that
  reaches a recv-site release with the hold set is the one whose own
  handleRecv just set it, with its SEND pending: releaseHoldSlow returns
  at its egress check without changing anything.
- A closing conn only has the flag cleared, which nothing observes
  (checkTimeouts skips closing conns; releaseConnState resets it).
Measured (86d46ac plus a counting overlay, -race, 8 MiB, one worker):
the recv-inline site was entered 60,559 times with the hold set and
changed nothing, processCQE's recv site 6 times, nothing; in the same
runs the send-inline site released 86 of 90 entries (the positive
control). The send-site releases and the checkTimeouts rescue stay.

Tests (engine/iouring): TestTransplantReapFailureIsNotRetried,
TestNoReapWithoutAsyncCancelFlags, TestReapSuppressedAfterFailedHandOff,
TestReapedRecvLeavesNoLinkOrBuffer, TestAsyncCancelProbeClassifies,
TestAsyncCancelProbeOnThisKernel, TestWorkersCarryTheAsyncCancelProbe
and TestReapOnTheRunningKernel; TestOneOwnerPerHandoff gains the split
and a closing-conn case; kernel_cancels_the_armed_recv also checks a
kernel that rejects the flags. The new EngineMetrics fields are summed
by the adaptive engine and pinned by both metrics tests.
…off loss (celeris#657)

TestFlapConnsPerRing fixes both engines at two workers, so 256
keep-alive connections are 128 per io_uring ring whatever memlock
funds, and runs three promote/revert cycles under continuous load,
requiring zero client errors. Step 0 validated it as the CI gate for
face 2 (main FAIL 6/6), and the #681 campaign measured it: base 5b2e83b
FAIL 8/8 (1,713 lost requests, each one a stale recv the witness
counted), fix PASS 8/8. Until now it existed only as an overlay.

The file is the measured step-0 source with four changes: its headers
no longer say "overlay only", the conditional six-cycle variant (never
triggered) is dropped, the two files are one, and a premise guard skips
it when RLIMIT_MEMLOCK funds fewer than two io_uring workers, failing
instead under CELERIS_REQUIRE_UPSWITCH=1 as the adaptive CI job sets it.
…r leg, and run the new fd-lifetime tests (celeris#657)

- adaptive race step (memlock unlimited): now shell: bash with
  pipefail and a tee, plus a tally for TestFlapConnsPerRing: exactly one
  top-level RUN and PASS, no SKIP line for it, and exactly one RESULT
  line showing both engines at two workers and err=0. At raised memlock
  the two revert tests did not lose on the base; T4 did in 8 of 8 runs.
- one-worker revert leg: the verdicts of TestReverseTransplant and
  TestBidirectionalFlap tolerate up to 64 lost requests (a base flap run
  PASSED with 56), so the step also requires each test's own summary
  line (async=false) to read err=0.
- unit witness step: the eight new fd-lifetime tests join the PR-2
  list, 29 in all.
…load line (celeris#657)

The celeris657 load line of the engine-level fd-lifetime tests now also
prints TransplantClaimDeferred, TransplantReapFailed and
TransplantReapUnsupported (read by name, -1 on a tree without them), so
a run on a kernel without cancel flags shows which path its connections
took.
… and pin that data lifts the reap suppression (celeris#657)

- probeAsyncCancelFlags is probeAsyncCancel(IORING_ASYNC_CANCEL_ALL).
  TestAsyncCancelProbeOnThisKernel also submits a cancel_flags bit no
  kernel defines, which every kernel rejects with -EINVAL, so a probe
  that stopped reading the kernel's answer fails on a 5.19+ kernel too.
- TestReapSuppressedAfterFailedHandOff now ends with the conn receiving
  data again, being served with the drain stopped, and having its linked
  RECV reaped once the drain is back: the suppression lasts until data,
  not forever.
@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 2 is pushed at 48dfe1a. Each finding of the two round-1 reviews is answered below. The PR body is rewritten to match; every figure there and here is regenerated by a script named next to it, over saved logs.

Correctness review

C1 (major), fixed. Verified first: I read the kernel source, and the verifier measured on real Ubuntu archive kernels under QEMU. On 5.15.0-191 the reap's form (CANCEL_ALL, reported) returns -EINVAL and leaves the recv armed; 5.19.0-50 and 6.8.0-138 accept it.

  • handleTransplantReap now retries only a miss (0 or -ENOENT). Any other non-hit result is counted in the new TransplantReapFailed (must stay 0) and is neither retried nor followed by a hand-off.
  • A startup probe, probeAsyncCancelFlags, submits the reap's exact SQE against a user_data nothing carries: res >= 0 or -ENOENT means supported, and -EINVAL means not. New stores and logs the answer, and createWorkers copies it to every worker.
  • Where the probe says unsupported, startReap places nothing and queues no retry, and counts TransplantReapUnsupported. The connection keeps its recv until the recv completes on its own. A sync connection is then HELD and leaves with nothing in flight; a promoted async connection stays on io_uring (placement only).
  • New tests: TestTransplantReapFailureIsNotRetried (the -EINVAL, -EBADF and -ECANCELED results) and TestNoReapWithoutAsyncCancelFlags (both sites). Both fail on 86d46ac: "loop iteration 1 placed [ASYNC_CANCEL …] after the reap failed with -22" and "the SEND completion placed [ASYNC_CANCEL …] on a kernel without cancel flags".
  • Also new: TestAsyncCancelProbe*, TestWorkersCarryTheAsyncCancelProbe and TestReapOnTheRunningKernel.
  • Mutants n01 to n10 are all killed, including the requeue restored, the probe ignored or never stored, and a probe that stops reading the kernel's answer.
  • On the verifier's rig, with the same kernel .deb hashes, this head's test binary on 5.15.0-191 gets the probe answer false, and the kernel reap subtest takes the rejected-flags branch: no retry, no hand-off. The same subtest from 86d46ac FAILs there with "loop iteration 1 placed 1 SQE(s) after the kernel rejected the reap". At engine level on 5.15, this head made 0 reaps against 86d46ac's 85 failing ones, with W1 = W2 = 0 and 0 client errors on both. The QEMU binaries are static, so there is no -race in those runs.
  • The pre-existing breakage (cancelConnOps, the WS pause, PauseAccept, the driver cancels, and EXT_ARG on 5.10.0-161) is not changed here. It is filed as io_uring: every async cancel fails with -EINVAL on Linux 5.10-5.18 (IORING_ASYNC_CANCEL flags are 5.19+) #682 (v1.6.0, bug + engine/iouring) with the measured facts and the sites at b79888e. The batch-truncation point (A5 parking with a failing cancel still queued) is in that issue too: REAP can no longer be the failing SQE.

C2 (minor), fixed. handOff sets reapSuppressed when the dup or SetNonblock fails, startReap places nothing while it is set, and handleRecv clears it when data arrives. The dup goes through a Worker.dupFD seam.

  • TestReapSuppressedAfterFailedHandOff forces EMFILE through the seam. The completion then places only the re-armed RECV, nothing on later iterations, one dup try per request, and a reap again once the connection has received data and the dup works.
  • On 86d46ac plus only that seam, the test fails with "placed [RECV … ASYNC_CANCEL …]".
  • Mutants n11 to n14 are killed: the suppression dropped, never set or never lifted, and the seam ignored.

C3 (minor), fixed, as the verifier proposed.

  • tryTransplant leaving a claimed async connection to its claim now counts TransplantClaimDeferred, a rate.
  • TransplantDoubleClaim counts only a connState that no longer owns its slot in finishAsyncTransplant, which in async mode is a second hand-off of the same connection. A closing connection and the range check are refused uncounted.
  • A gate of TransplantDoubleClaim == 0 (DECISION step 5, residual 7) now means what it says. engine.go's docs, the PR table and the gate list are updated.
  • TestOneOwnerPerHandoff pins the split, plus a closing-connection case, and fails on 86d46ac with "TransplantDoubleClaim = 1, want 0". Mutants n17 to n19 and n22 are killed.
  • Not done: the optional clearing of transplantPending in drainDetachQueue's detachClosed and asyncClosed branches. Those entries are connections being closed, releaseConnState resets the flag, and with the claim now counted as a rate a stale one only adds to that rate. It would be a new write on the close path for no gate.

C4 (nit), fixed. reapOutcome returns a flagged provided buffer and clears recvLinked. The verifier found no consequence of either omission, so this is defensive. TestReapedRecvLeavesNoLinkOrBuffer fails on 86d46ac ("recvLinked is still set", buffer not returned). Mutants n15 and n16 are killed.

C5 (nit), acknowledged, unchanged. The close paths' fd-number hazard is pre-existing and not touched here. DECISION step 5 files it with a milestone after the fix lands, as planned. On 5.10 to 5.18 the same close paths also send a cancel the kernel rejects; that part is in #682.

C6 (nit), fixed. The A5 comment now says an idle SQ thread needs its NEED_WAKEUP kick, and that no tier enables SQPOLL. I changed the comment and no code, because no test could reach an SQPOLL branch.

C7 (nit), disclosed. resource.MinWorkers is 2, so Resources.Workers: 1 variants cannot be built. The one-worker legs rely on the runner's 8 MiB memlock, and T4 pins two workers explicitly. See the PR body's "Departures".

Evidence review

E1 (major), fixed. T4 is committed as adaptive/flap_conns_per_ring_test.go.

  • It is the measured step-0 source with four listed changes: headers, the six-cycle variant dropped, one file, and a premise guard that skips below 24 MiB and fails under CELERIS_REQUIRE_UPSWITCH=1.
  • The adaptive race step (memlock raised) now tallies it: exactly one RUN and one PASS, no SKIP line for it, and exactly one S0T4 RESULT line with conns=256 wE=2 wI=2 conns_per_ring=128 … err=0.
  • Proven both ways with the step's verbatim run: block. The committed tree passes, with the whole package at 76 PASS, 0 FAIL and 0 SKIP. These arms all fail:
    • T4 deleted, renamed, t.Skip'd, or with a skipped subtest;
    • io_uring refused by seccomp;
    • T4 at three workers (the test passes; the RESULT line is refused);
    • the no-HOLD/REAP engine (T4 FAILs with err=141);
    • that engine with T4's verdict neutralised (go test ok; the step fails on err=250 alone).
  • One thing the seccomp arm surfaced: with io_uring unavailable, 16 adaptive tests skip and go test is ok, so before this change that step went green. T4's SKIP now turns it red.

E2 (minor), fixed. The one-worker leg now also requires each test's own summary line ([async=false] … err=N) to read err=0.

  • The discriminating arm neutralises the two scenarios' verdicts (every t.Errorf becomes t.Logf) on the no-HOLD/REAP engine. go test passes, and the step fails on the summaries alone (err 0 and 50).
  • The same neutralisation on this head's engine passes (err 0 and 0).
  • The plain no-HOLD/REAP arm fails too; there the verdicts catch it (err 14 and 44).

E3 (minor), fixed. tools/validity_r2.py reads the celeris657 engine workers=N witness as well as the listening line, and voids an engine-level unit arm that prints neither. Re-run over all 308 round-1 B, C and S logs (tools/e3_rerun.py, tooling/E3-RERUN.txt):

  • 0 verdicts change;
  • the 26 unit arms that build an engine are now actually checked, and all read workers=1; 10 build no engine;
  • regate's worker check, applied as a rule, finds 0 violations.

E4 (minor), fixed. The summary now says 0 keep-alive failures and 0 churn read timeouts on the switch-stall harness. The churn census (fix: reset 189, refused 16, dial reset 3) is in the body, from verify/tools/v7_e4_churn_census.py.

E5 (nit), fixed.

  • "438 container logs plus a build log" (verify/tools/v7_e5_logcount.sh).
  • The seccomp arm's FAIL count is now printed by the verdict tool: 22 top-level, 53 with subtests.
  • The switch-stall W1 = keep-alive failures + churn read timeouts equality is attributed to verify/tools/v7_e5_w1_join.py, which computes it.
  • The mutation verdict prints each run's worker witness, and says "no engine built" for the adaptive-metrics packages.
  • The CI summary counts worker witnesses only from engine and test log lines, so GitHub's echo of a step's grep pattern can no longer inflate it (tools/summarize_ci_r2.py).

E6 (nit), fixed. Each of the three steps has a skipped-subtest arm, and all three fail on the indented SKIP line. The concurrency sampler now polls every 0.2 s from launch: every arm saw exactly one container, including the seccomp arms (2 samples for the unit one).

E7 (nit), fixed. The body discloses the Workers: 1 departure. It also says that raising the runner's memlock turns the witness step and the one-worker leg red by design.

Also open from round 1

m15, removed as equivalent. The verifier proved it, from code and by a counting overlay: 60,559 recv-site entries with the hold set, none changing anything, and the send site releasing 86 of 90 in the same runs. The call is removed at both recv dispatch sites; processCQE's recv site had the same status, because m12 was killed only through the send site. The proof is in commit 667d4d3's message. Round 2's mutation run: 54 mutants, 54 killed, 0 survivors (mut/VERDICT.txt).

Refused dials at the revert (post hoc): still open, and not changed by this PR. The verifier read it from code:

  • performSwitch is byte-identical in both trees, and nothing under engine/epoll changed.
  • HOLD, REAP and A5 are inert while the listener is handed over: the drain starts only after PauseAccept returns, and A5 needs listenFD < 0.
  • The gap is the synchronous io_uring listener close against epoll re-creating its listener at a loop's next epoll_wait return.

The falsifiable hypothesis and the measurement it needs (per-revert in-process timestamps, a 1 ms LISTEN sampler, stratified by suspended epoll loops) go to PR 3, which owns the switch.

Suites, lint and CI for this head are in the body. In short: 0 items go from PASS on 86d46ac to anything else on this head, across 10 whole-package cells; lint is clean in all five modules; and CI run 35449608095 had 11 of 11 jobs green on the first attempt, with its tally lines quoted in the body.

@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 2: refused dials at a revert, measured (pre-registered)

Verdict: no_regression (F - B -0.0192 <= A/A floor 0.0577 (Fisher p 0.7676)). This comment was generated by tools/mkcomment.py from results.json, which tools/analyze.py M writes from the raw logs. The evidence is under evidence/celeris-657/impl/pr2-r2/measure/.

The design was registered in 00-PREREG.md before the first build: sha256 ca4545bde56e85f458398b23c701e4fcf2b8452708627fa5dff309a6c4a48172, hashed 2026-09-19T15:17:47Z UTC. Three amendments were made before any campaign run (A1-A3 in AMENDMENTS.txt):

  • A1: the listener sampler moved to a separate C process, because the in-process Go sampler ran at about 1.7 ms. The new sampler has a median gap of about 1.0 ms.
  • A2: analyzer fixes found on the excluded smoke runs.
  • A3: a wrapper script.
    One amendment came after the data: A4 adds the post-hoc descriptive section below. It changes no registered outcome, prediction or verdict.
  • Arms: base = main b79888e, fix = PR head 48dfe1a, and same-bytes A/A halves of each (Ba/Bb, Fa/Fb).
  • Size: 104 runs per arm (exact Fisher power 0.839 for 1/80 vs 9/80) in 52 counterbalanced Williams blocks, seed 6572019. Each block carries one control run: 12 POS (every epoll re-listen delayed 20 ms) and 40 NEG (ResumeAccept waits for an epoll listener).
  • Shape: round 1's SW shape (5k conn/s churn, 32 keep-alives, P1/R1/P2/R2), 8 MiB memlock, so io_uring runs 1 worker and epoll 4 loops. The workers= census voids any other count.

Primary: runs with a refused churn dial within ±100 ms of a revert

arm runs positive 95% CI
base (Ba+Bb) 7/104 (0.067) [0.027, 0.134]
fix (Fa+Fb) 5/104 (0.048) [0.016, 0.109]
Ba 3/52 (0.058)
Bb 4/52 (0.077)
Fa 4/52 (0.077)
Fb 1/52 (0.019)
  • Same-bytes A/A: |Ba−Bb| = 0.0192 (Fisher p 1.000); |Fa−Fb| = 0.0577 (p 0.363). The A/A floor is 0.0577.
  • fix − base = -0.0192, and per half Fa−Ba +0.0192, Fb−Bb -0.0577. Two-sided Fisher on fix vs base gives p = 0.7676.
  • Refused dials per revert: base 8 over 208 reverts, fix 13 over 208. Reverts with ≥1 refusal: base 7/208 (0.034), fix 5/208 (0.024). Per-run Mann-Whitney p = 0.6038.

Stratified by whether any epoll loop was parked at ResumeAccept

stratum base refused reverts fix refused reverts Fisher p
S1 (≥1 loop parked) 7/208 (0.034) (8 dials) 5/208 (0.024) (13 dials) 0.7711
S0 (no loop parked) 0/0 (0 dials) 0/0 (0 dials) nan
placement: share of reverts in S0 0/208 (0.000) 0/208 (0.000) 1.0000

Instrument

  • POS: the injection fired in 24/24 reverts, and 24/24 reverts refused.
  • NEG: the injection fired in 80/80 reverts, with 0 refused dials.
  • IV gate: PASS.
  • Witness (engine trace + 1 ms netlink LISTEN sampler):
    • refused reverts where an epoll listener was already listening before io_uring's close: 0;
    • S1 reverts with a listener-free gap: 19;
    • S1 reverts that refused: 12;
    • refused dials inside the witnessed gap: 21/21.

What the witness shows (post hoc, descriptive: A4)

  • At every revert of both arms, at least one epoll loop was parked at ResumeAccept: stratum S1 held all 208 base and 208 fix reverts. So the no-loop-parked case that H-R names never occurred, and the fix did not change where connections sit (placement p = 1.000).
  • The hand-over is a race between two asynchronous steps. ResumeAccept wakes the parked loops and returns without waiting for any of them to listen. Then io_uring's PauseAccept closes the one io_uring listener.
    • base: the first epoll listen() returned a median 67 µs after the resume call (p95 158, max 254). io_uring's close returned a median 125 µs after the pause call (p95 703). Outcomes: 8 reverts with a listener-free gap (G=pos, gaps 15-91 µs), 7 too close to call (amb), 193 with an epoll listener in LISTEN first (neg). Refused reverts by G: {'pos': 5, 'amb': 2}.
    • fix: the first epoll listen() returned a median 66 µs after the resume call (p95 145, max 304). io_uring's close returned a median 115 µs after the pause call (p95 637). Outcomes: 11 reverts with a listener-free gap (G=pos, gaps 21-109 µs), 7 too close to call (amb), 190 with an epoll listener in LISTEN first (neg). Refused reverts by G: {'pos': 5}.
  • G=pos, fix vs base: Fisher p = 0.640.
  • Every refused dial ended inside the witnessed gap (21/21). No refused revert had an epoll listener up before io_uring's close.
  • NEG makes ResumeAccept wait for an epoll listener. It had 0/80 refused reverts, against 12/416 for base+fix (Fisher p = 0.229, post hoc).
  • Reading: the refusals are a race that already exists on main. The PR does not touch either end of it (epoll ResumeAccept, io_uring's pause close). It is not a regression of this PR.

Registered predictions

  • Q1 HELD: every scheduled run valid at its first attempt. Measured: 260/260.
  • Q2 HELD: POS: injection fired in every POS revert; >= 1 refused dial in >= 22 of 24 POS reverts; G=pos in >= 22. Measured: fired 24/24, refused 24, G=pos 24.
  • Q3 HELD: NEG: injection fired (neg_wait ok and G=neg) in every NEG revert, and 0 refused dials in every NEG revert window. Measured: fired 80/80, refused dials 0.
  • Q4 HELD: H-R (a): no refused revert in any arm has G=neg. Measured: 0 [].
  • Q5 MISSED: H-R: no B or F revert in stratum S1 has G=pos. Measured: 19.
  • Q6 MISSED: H-R: no B or F revert in stratum S1 has a refused dial. Measured: 12.
  • Q7 HELD: H-R (b): within S0 and within S1, F vs B refused-revert Fisher p > 0.05. Measured: S0 p=nan, S1 p=0.7711.
  • Q8 HELD: >= 90% of refused dials in B/F revert windows end inside [close call - 1 ms, first epoll listen + 1 ms]. Measured: 21/21.
  • Q9 HELD: primary: F - B <= A/A floor and Fisher p > 0.05. Measured: F-B -0.0192, floor 0.0577, p 0.7676.
  • Q10 MISSED: base primary rate <= 5/104. Measured: 7/104.
  • Q11 HELD: refused dials per run (revert windows): F vs B Mann-Whitney p > 0.05. Measured: p 0.6038.
  • Q12 HELD: sampler: median inter-sample gap within +-100 ms of B/F reverts <= 1.5 ms; POS zero-listener sample in >= 20/24 reverts; NEG zero-listener samples in 0 reverts. Measured: median 0.997 ms, POS 24, NEG 0.
  • Q13 HELD: no phantom: io_uring's listener absent from every sample starting >= 1 ms after its close returned. Measured: 0 reverts.
  • Q14 HELD: fix non-vacuity: TransplantHeld+TransplantReaps > 0 at the end of every F run, and F keep-alive request failures lower than B's (Mann-Whitney p <= 0.05). Measured: 104/104; ka failures B 59 F 0 p 6.206e-05.
  • Q15 HELD: no DATA RACE, panic, FAIL or SKIP in any log. Measured: 0 runs.
  • Q16 HELD: 0 refused dials outside the revert windows in every B and F run. Measured: 0.

Scripts (all under measure/tools/):

  • power.py: n.
  • mksched.py: the schedule.
  • mkovl.py: the overlays.
  • build.sh, run1.sh and campaign.sh: build and run.
  • validity.py: void rules.
  • analyze.py: REPORT.txt and results.json.
  • mkcomment.py: this comment.

How much this rules out (added by the maintainer after review)

  • Power. The registered power (0.839) assumed round 1's base rate of about 1/80. The observed base rate was 7/104 (Q10 missed), so the power against round 1's alternative (9/80 on the fix) is lower than registered.
  • Tolerance of the rule. The registered verdict rule would also have called a fix increase of up to +0.0577 (the A/A floor) no_regression.
  • What the conclusion rests on. The fix's point estimate is below base (5/104 vs 7/104). More importantly, the witness puts every one of the 21 refused dials inside a listener gap that opens on both trees for the same reason (ResumeAccept does not wait for an epoll listener), and neither end of that gap is touched by this PR. Round 1's 9/80 vs 1/80 did not reproduce.
  • The race itself. It is a real defect on main, independent of this PR, and is now tracked in its own issue. The NEG control (ResumeAccept waits for an epoll listener) removed every refusal, 0 of 80 reverts.

…ot reap (celeris#657)

Review R1 on #681. Where the kernel rejects IORING_ASYNC_CANCEL flags, a
promoted async conn under a reverse drain never left io_uring, yet its
dispatch goroutine claimed the hand-off at every park: the worker's drain
of the claim found the recv the feed path had armed, which only a reap can
clear, and refused; the next request respawned the goroutine. A spawn and a
detach-queue round trip per request, with no end while the drain lasted.

asyncTransplantEligible now refuses on a worker without cancel flags, so the
goroutine parks as it does with no drain set and the conn stays on io_uring
(placement only). asyncCancelFlags is set before the worker starts and never
written again, so the dispatch goroutine reads it race-free.

Tests: TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing runs the
real dispatch goroutine through three requests with the drain set and
requires no claim and one goroutine for all three; on 48dfe1a it fails with
3 of 3 parks claimed, 3 goroutines and TransplantReapUnsupported = 3. Its
control (control_async_park_claims_with_flags) shows the same rig sees a
claim, a reap and the hand-off where the worker can reap.
TestReapSuppressedAfterFailedHandOff/async_conn_retries_at_its_next_park pins
the async side of a failed dup: the data that respawns the goroutine clears
reapSuppressed before any park, so the hand-off is retried at the next park,
once per request, and the conn leaves once the dup works.
…rejection (celeris#657)

Review R2 on #681. probeAsyncCancelFlags cached a probe that never reached
the kernel (a ring that could not be set up, a failed submit or wait, a
missing or foreign completion) exactly as it cached -EINVAL, and New logged
both at Info as "probe failed". A rejection is the kernel's answer; a probe
with no answer says nothing about the kernel, and on 5.19 or later it turns
the hand-off's reap off where it should work.

The probe now returns asyncCancelAccepted, asyncCancelRejected (-EINVAL, or
a result it does not recognise) or asyncCancelNoAnswer. Only accepted turns
the reap on. logAsyncCancelProbe logs a rejection at Info and no answer at
Warn when the kernel's version is 5.19 or later (Info below it), each with
the reason and the kernel; the engine-selected line also carries
async_cancel_probe. The probe's ring is made through newAsyncCancelProbeRing,
a variable, so a test can make the probe fail before the kernel answers.

The docs that said the next response is then HELD now state HOLD's
condition: a worker without a provided-buffer ring. A kernel that rejects
the flags has none (buffer rings arrived in 5.19 too); where one exists
because the probe got no answer, a sync conn is not held and stays on
io_uring.

Tests: TestWorkersCarryTheAsyncCancelProbe/no_answer_is_logged_apart_from_a_rejection
runs New with the probe's ring failing: on 48dfe1a plus only the ring seam it
fails, the record at INFO reading "async cancel flags runtime probe failed",
want WARN. TestAsyncCancelProbeClassifies gains no_answer_is_not_a_rejection
and log_levels (the level matrix).
@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 3 is pushed at 272bcba. It answers R1 to R8 from the two round-2 re-reviews. One half of R1 is disputed, with the evidence below; everything else is fixed. The PR body is updated to match. Every figure here and there comes from a script named next to it, run over saved logs. Round 3's evidence is under evidence/celeris-657/impl/pr2-r3 (manifest SHA256SUMS, index INDEX.txt).

Commits: 16c9719 (R1), 5e9c7f0 (R2), c15c686 (R3), 356fbe2 and 1037b66 (R4), c173711 (R5), and 1f883e2, a test-only fix for two races in round 3's own async rig, found by the local interlock run and a stress run. After those come 683c7bb (the deliberate red commit, R6) and 272bcba, its revert. The final tree is byte-identical to 1f883e2's.

R1 (minor): fixed where the worker cannot reap. The reapSuppressed half is disputed.

  • Fixed: asyncTransplantEligible refuses when !w.asyncCancelFlags, so on such a worker the dispatch goroutine parks as it does with no drain set. The field is written before the worker starts and never again, so the goroutine's read is race-free.
  • TestNoReapWithoutAsyncCancelFlags/async_park_claims_nothing runs the real dispatch goroutine through three requests with the drain set on a worker without flags. It requires no claim, one goroutine for all three requests and nothing placed but the feed path's RECVs.
    • On 48dfe1a it fails: "the dispatch goroutine claimed the hand-off at 3 of 3 parks, and the 3 requests ran on 3 goroutines, want 0 claims and 1 goroutine" and "TransplantReapUnsupported = 3, want 0" (ff/VERDICT.txt).
    • Its control, control_async_park_claims_with_flags, shows the same rig sees the claim, the reap and the hand-off where the worker can reap.
    • Mutants r01 (the check removed) and r02 (the check inverted) are killed.
  • Disputed: "or the conn is reapSuppressed". That condition cannot hold at a park, so a check of it there would be dead code. At 48dfe1a:
    • handleRecv clears the flag on every data completion (worker.go:2557), before the async feed that respawns the goroutine (worker.go:2716-2751) or wakes it. The inline promotion's spawn (worker.go:4019) also comes after that clear.
    • A promoted connection is handed off only while its goroutine is not running. tryTransplant returns while it runs (transplant_source.go:94-97), rerunHandOff checks the same (fd_lifetime.go:185-191), and drainDetachQueue acts only on a claim the goroutine made on its way out (worker.go:4100-4103, 4556-4557).
    • So the flag is set after the goroutine exits and cleared before it next runs. It is false at every park.
    • Reading it from the dispatch goroutine would also race that unconditional worker-side write.
  • What actually happens after a failed dup: the next park claims, and the claim costs one reap and one dup per request while the dup keeps failing. The connection leaves at the first park after the dup works. A sync connection pays one dup per request at its held response, the same cadence.
    • TestReapSuppressedAfterFailedHandOff/async_conn_retries_at_its_next_park pins this. It passes on 48dfe1a as well, since nothing there changed.
    • Mutant r03 is the "sticky" alternative that a park-side check would need: data no longer clears the flag for a promoted connection, and the park refuses while it is set. The pin kills it: after one failed dup the connection is stranded ("parked without claiming").
  • So R1's per-request cost with no end is the no-flags case, and that case is fixed. Both halves are documented at startReap and on the field (see R5).

R2 (nit): fixed.

  • The probe now returns accepted, rejected (-EINVAL, or a result it does not recognise) or no answer (ring setup, submit or wait failed, no completion, a foreign completion). Only accepted turns the reap on.
  • logAsyncCancelProbe logs a rejection at Info. It logs no answer at Warn on a kernel whose version is 5.19 or later, and at Info below that. Each record carries the reason and the kernel, and the engine-selected line carries async_cancel_probe.
  • The probe's ring comes from newAsyncCancelProbeRing, a variable, so a test can make the probe fail before the kernel answers.
  • TestWorkersCarryTheAsyncCancelProbe/no_answer_is_logged_apart_from_a_rejection runs New with that ring failing. On 48dfe1a plus only the ring seam (ff/ovl/FF-48-R2SEAM/probe.go.seam.diff) it fails, logged "at INFO as "async cancel flags runtime probe failed: …", want WARN".
  • TestAsyncCancelProbeClassifies gains no_answer_is_not_a_rejection and log_levels. Mutants r04 to r10 are killed.
  • The HOLD docs (engine.go at the probe, and fd_lifetime.go's header and startReap) now state the condition: no provided-buffer ring on the worker. A kernel that rejects the flags has none, because buffer rings also arrived in 5.19. Where one exists because the probe got no answer, a sync connection is not held and stays on io_uring.

R3 (nit): fixed.

  • The kernel version is only logged now (kernelRelease), and an unparsable release no longer fails anything. The bit-31 rejection check stays as the proof that the probe reads the kernel.
  • What the version check used to catch, a probe that answers wrongly on this kernel, is now checked against the kernel itself. TestReapOnTheRunningKernel/probe_matches_the_kernel forces the flags on and lets the running kernel answer the real reap, which either cancels the recv or fails it with -EINVAL. It then requires the probe's answer to match.
  • Mutant r11 (a probe that always says rejected) is killed by it with "the probe's answer is not what this kernel does". n08 and n09 (a probe that stops reading the kernel) are still killed by the bit-31 check.

R4 (nit): fixed.

  • T4 now also asserts, from e.Metrics(): W1 = 0, W2 = 0, and TransplantDoubleClaim, TransplantHoldRescued and TransplantReapFailed = 0, each with its own S0T4 …: message.
  • It also asserts that every switch moved the connections. The engine switched away from must have detached at least 256 and hold none, and the one switched to must have adopted at least 256 and hold all of them. All 54 switches of this round's 9 unmutated local T4 runs moved and adopted 256 (t4-switch-census.txt, tools/t4_switch_census.py), and so did the 6 of the green CI run.
  • The header is corrected.
  • Its io_uring-availability and adaptive.New skips now fail under CELERIS_REQUIRE_UPSWITCH=1, like its memlock skip. With io_uring refused by seccomp, 48dfe1a's T4 SKIPs and this head's FAILs "… forbids skipping" (ff/VERDICT.txt). The rest of that blind spot is CI: with io_uring unavailable, 16 ./adaptive tests skip and the adaptive job passes even under CELERIS_REQUIRE_UPSWITCH=1 #684, as asked.
  • The RESULT line gains the new fields after err=, so ci.yml's tally is unchanged.
  • The engine-level load helper fails on any non-zero TransplantDoubleClaim, TransplantReapFailed or TransplantHoldRescued, for every caller. drain_stops_while_held checks the first two as well.
  • Each new assertion is shown to fire on a planted defect: mutants r12 to r20 (below). r20 turns the reverse drain off. The connections then stay on io_uring and nothing is lost, so the MOVE check is the only thing that catches it.

R5 (nit): fixed (docs only).

  • TransplantReapUnsupported now says it counts only connections the worker serves itself, and that those are held only without a provided-buffer ring. It also says a promoted async connection is never offered for the hand-off there, stays on io_uring and is not counted.
  • reapSuppressed's doc now says only data and the release clear it, and that it outlives its drain. It also says what that does across drains: placement only, and only at a completion that is not data, such as an -ENOBUFS re-arm. I documented it rather than reset it at StartTransplant: that runs on another goroutine, and the flag is worker-owned.

R6 (nit): done. Run 35465814564 is red on purpose.

  • Its commit, 683c7bb ("DO NOT MERGE: plant the no-HOLD/REAP engine and neutralise the T4 and revert-test verdicts …"), plants round 2's no-HOLD/REAP engine. It also turns every t.Errorf in T4 and in the two revert scenarios into t.Logf, so the adaptive steps can go red only through their own tallies. That is the check R6 asked to see on the runner. With the verdicts live, both steps stop at go test before their tallies run (round 2's T-NHR and A-NHR arms).
  • From the real job logs (ci/683c7bb/SUMMARY.txt):
    • Race step: go test ok, 76 PASS. celeris#657 T4: … RESULT lines 1, at two workers with err=0: 0, with RESULT err=188 w1=188 w2=398 … held=0 reaps=0. T4's own messages still print, as logs: LOSS, W1, W2, and MOVE at the five switches after the losses.
    • One-worker leg: go test ok, 2 of 2 PASS. … summaries 2 with err=0: 0, err 13 and 36.
    • The unit job went red too (its root race step: the engine-level tests on that engine), and so did lint (unused: startReap, onlyRecvInFlight, holdEligible, noteReap and noteReapUnsupported lost their callers). Every other job passed.
  • Rehearsed locally first from the same verbatim steps (red/rehearsal-1f883e2.out: err=249; summaries err 23 and 52).
  • Run 35465988505 on the revert, 272bcba, is green. 272bcba's tree is identical to 1f883e2's (tree 2759f73).
  • Both runs, and what went red, are in RED-RUNS.txt.

R7 (nit): fixed. P4's figures (60,559, 6, and 86 of 90) now cite verify/tools/v4_tally.py → verify/v4/V4-TALLY.txt. The script was re-run for this round, and its output is identical (pr/V4-TALLY-regenerated.txt).

R8: fixed. Residual 5 is rewritten around the measured result: no_regression (fix − base = −0.019 within the A/A floor of 0.058, Fisher p = 0.77), a pre-existing ResumeAccept/listener-close race now tracked as #683, and a link to the measurement comment. A "Tracked elsewhere" section lists #682, #683, #684 and #685.

Checks on this head.

  • Failing first (ff/VERDICT.txt, tools/ff_verdict_r3.py): R1's test fails on 48dfe1a plus only the head's two fd-lifetime test files. R2's fails on 48dfe1a plus only the ring seam. T4 SKIPs on 48dfe1a and FAILs here with io_uring refused under CELERIS_REQUIRE_UPSWITCH=1. All of them pass here.
  • ci.yml is unchanged. The new tests are subtests of the 29 the witness step names. The interlock arms were re-run on this head from the steps' verbatim run: blocks: W-GOOD, A-GOOD, T-GOOD and T-BLIND-GOOD pass; T-SECCOMP, T-WORKERS3, T-NHR and T-BLIND-NHR fail for their planted reasons. That is 8 of 8 arms (interlock/1f883e2/VERDICT.txt, tools/il_verdict_r3.py). T-SECCOMP's reason changed: T4 now fails instead of skipping.
  • Mutation (mut/VERDICT.txt, tools/analyze_mut_r3.py): 74 mutants, round 2's 54 plus r01 to r20. 73 are killed by their named tests, for their registered reasons.
    • One is not killed by all its named tests: m04. It is killed by TestTransplantNeverHandsOffArmedRecv and TestHandoffHasNothingInFlight, but TestStaleRecvDataCounted passed. In that run the test never executed the mutated branch: W2=0, and all 128 connections were held.
    • How many connections reach the reap in that test is scheduling-dependent. On the unmutated code it ranges from 0 to 88 across 10 runs, and is 0 in 2 of them (mut/M04-EXPOSURE.txt, tools/explain_m04_r3.py). Round 2's m04 run had 40, and the test failed there.
    • I report it as a finding about that test, which is unchanged since round 1. It is not a gap in this head: the branch's deterministic coverage is the other two tests.
  • Suites, this head against 48dfe1a (suites/COMPARE.txt, tools/compare_suites_r3.py), -race -v, 10 cells: ./engine/iouring at 8 and 128 MiB, ./adaptive at 8 MiB with the CI skip and at 128 MiB with REQUIRE_UPSWITCH, ./engine/epoll, root, and the websocket CI step.
    • 0 items go from PASS to non-PASS. There are no FAIL, race or panic lines in any run.
    • The only differences are the 7 new subtests.
  • Lint as the lint job runs it is clean (lint/, tools/run_lint.sh).
  • CI on 272bcba: runs 35465988505 and 35465987011, all 11 jobs green at the first attempt, none re-run.
  • One more thing this round found in itself: the first local interlock run of round 3's own async rig failed once ("the drain of the claim placed []"), because the rig counted a claim before it was queued. A stress run also caught poll(2) returning EINTR. Both are races in the test, fixed in 1f883e2. After the fix, 1000 of 1000 runs of the two tests passed under -race (dev/stress-r1b.log), and 290 of 290 of the witness list (dev/stress-witness.log). The failing run is kept (interlock/superseded-c173711-W-GOOD-flake/).

…er (celeris#681 N3)

The zero value of asyncCancelProbe was asyncCancelAccepted, the one
answer that turns the hand-off's reap on, so an answer that was never
set read as ON. The zero value is now asyncCancelNoAnswer, which keeps
the reap off. TestAsyncCancelProbeClassifies/zero_value_is_no_answer
pins it.

asyncTransplantEligible's doc now says that a probe with no answer has
the same effect there as a rejection.
…its own class, and warn with the errno (celeris#681 N2)

classifyAsyncCancelProbe read any completion other than an acceptance
(res >= 0, -ENOENT) or -EINVAL as a rejection, and New logged it at
Info, even on a 5.19+ kernel where the flags exist. It is now its own
class, asyncCancelUnexpected: the reap stays off, the reason names the
errno (for example "cqe.res=-9 (EBADF)"), and New logs it at Warn on
every kernel.

TestAsyncCancelProbeClassifies/unrecognised_answer_is_its_own_class
pins the class, the errno and the Warn; the classification table and
log_levels gain the new class. The docs that list what keeps the reap
off name it.
…e, and wait again after a cut-short wait (celeris#681 N1)

probeAsyncCancelFlagsCached kept whatever the probe returned for the
life of the process, a probe that got no answer included. One transient
failure, such as EMFILE or ENOMEM setting up the probe's ring while the
adaptive engine builds its io_uring engine under load, kept the
hand-off's reap off from then on. Now only the kernel's answer, accepted
or rejected, is cached. After no answer, or an answer the probe does not
recognise, the next New probes again. Calls are serialised by a mutex.

The probe waits for its cancel's completion only when it is not there
after the submit. SubmitAndWaitTimeout returns nil both when that wait
times out and when a signal cuts it short (EINTR), so the probe used to
read an interrupted wait as no answer. A wait that returns early with
nothing to read is now repeated once, for the rest of the 500 ms.

Test seams: runAsyncCancelProbe (the probe the cache runs), and the
probe's submit, wait and wait timeout. New subtests:
TestWorkersCarryTheAsyncCancelProbe/only_an_answer_is_cached and
/unexpected_answer_keeps_the_reap_off_and_is_not_cached, and
TestAsyncCancelProbeOnThisKernel/a_cut_short_wait_is_retried_once.
…celeris#681 N4)

The branch for a kernel that rejects the cancel flags now logs what it
saw (TransplantReapFailed, TransplantHeld, adopted), so a run on such a
kernel shows the branch ran. Test-only.
…id not find the cancel flags (celeris#681 N4)

Running the witness list on Ubuntu 5.15.0-191 under QEMU (N4) found
TestHandoffHasNothingInFlight/async failing there on round 3's head and
on this one alike: "0 of 128 busy conns were handed off, want all".
Without the cancel flags a promoted async conn cannot be reaped, so it
is never offered for the hand-off and stays on io_uring
(asyncTransplantEligible, celeris#681 R1). The subtest wanted every conn
handed off on every kernel.

Where the engine's probe did not find the flags, the async subtest now
wants none of the promoted conns handed off, with the loss checks
unchanged (no client error, W1 = W2 = 0). A new case,
async_without_cancel_flags, takes the probe's answer as a rejection
(runAsyncCancelProbe), so that branch runs on every kernel, CI's
included. Test-only.
@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 4 is pushed at 5ce4986. It answers N1 to N6 from the two round-3 re-reviews. The PR body is updated to match, and its first line now says the PR must be squash-merged (N6). Every figure here and there comes from a script named next to it, run over saved logs; round 4's evidence is under evidence/celeris-657/impl/pr2-r4 (manifest SHA256SUMS, index INDEX.txt).

Commits: 30f9f70 (N3), 0c01c1c (N2), 0265e82 (N1), and two test-only ones from the QEMU runs N4 asked for: 92a8170 (the rejecting branch logs what it saw) and 5ce4986 (an engine-level test that could not pass on a kernel without the cancel flags — see N4).

N1 (nit): fixed. Only the kernel's answer is cached, and a wait a signal cut short is repeated once.

  • probeAsyncCancelFlagsCached keeps accepted and rejected for the process. A probe that got no answer is not kept, and neither is an answer the probe does not recognise (N2): the next New probes again, and logs again. The probes are serialised by a mutex, so concurrent News run one at a time and never overwrite an answer. One transient NewRing failure — EMFILE or ENOMEM while the adaptive engine builds its io_uring engine under load — no longer keeps the reap off for the life of the process.
  • The probe waits for its cancel's completion only when the completion is not already there. SubmitAndWaitTimeout returns nil both when its timeout expires and when a signal cuts the wait short (EINTR, ring.go), so a wait that returns early with nothing to read was cut short: it is repeated once, for the rest of the 500 ms. A second early return, or a wait that ran its full time, is no answer, and the reason says how many waits it took.
  • New checks, each a subtest of a test the CI witness step names (so the runner runs them with skipping forbidden):
    • TestWorkersCarryTheAsyncCancelProbe/only_an_answer_is_cached: three calls with each class of answer, counting the probes (1, 1, 3), and then two News across a no-answer, the second of which must turn the reap on.
    • TestAsyncCancelProbeOnThisKernel/a_cut_short_wait_is_retried_once: the cancel is held back at the submit and the wait is cut short once (the kernel's answer after 2 waits), twice (no answer after 2), and once at full length (no answer after 1).
  • Failing first on 272bcba plus only round 4's test seams (ff/ovl/FF-272-SEAM/probe.go.seam.diff; ff/VERDICT.txt, tools/ff_verdict_r4.py):
    • "three calls with the probe answering no answer ran it 1 time(s), want 3: only the kernel's answer (accepted or rejected) is kept …", and "the first New's probe got no answer, the second's is accepted: asyncCancelFlags false then false, probes run 1; want false, true and 2".
    • "wait cut short once: the probe gave (no answer, "no CQE produced for the cancel (waited 500ms)") after 1 wait(s), want accepted after 2".
    • The seam patch alone leaves 272bcba's own probe tests passing (arm FF-272-SEAMCTL), so the seams change no behaviour.
  • Mutants killed: q01 (cache every answer), q02 (cache the unrecognised one), q03 (cache none), q04 (no retry), q05 (retry after a full-length wait), q06 (retry twice).

N2 (nit): fixed. A completion the probe does not recognise is its own class.

  • classifyAsyncCancelProbe returns asyncCancelUnexpected for anything that is neither an acceptance (res >= 0, -ENOENT) nor -EINVAL, with a reason that names the errno ("the cancel completed with cqe.res=-9 (EBADF), which the probe does not recognise"). It keeps the reap off, it is not cached, and New logs it at Warn on every kernel, not Info.
  • TestAsyncCancelProbeClassifies/unrecognised_answer_is_its_own_class pins the class, the errno and the Warn for -EBADF, -ECANCELED and -EPERM; the classification table and log_levels carry the new class; TestWorkersCarryTheAsyncCancelProbe/unexpected_answer_keeps_the_reap_off_and_is_not_cached runs two News on such an answer and requires the reap off, two probes and one WARN record each naming EBADF.
  • Failing first on 272bcba: "classifyAsyncCancelProbe(-9) = rejected ("… cqe.res=-9, which the probe does not recognise"), want a class of its own, "unexpected"", "reason … does not name the errno EBADF", and the log record at INFO where WARN is wanted.
  • Mutants killed: q07 (folded back into rejected), q08 (logged at Info), q09 (New turning the reap on for it), q11 (the reason without the errno).

N3 (nit): fixed. The zero value of asyncCancelProbe is asyncCancelNoAnswer, so an answer never set keeps the reap off. TestAsyncCancelProbeClassifies/zero_value_is_no_answer pins it; on 272bcba it fails with "the zero value of asyncCancelProbe reads as accepted: an answer that was never set would turn the hand-off's reap on". Mutant q10 (the old order) is killed. asyncTransplantEligible's doc now says a probe with no answer, or one it does not recognise, has the same effect there as a rejection; the other places that list what keeps the reap off say it too (engine.go's field and New, worker.go's field, startReap, and TransplantReapUnsupported in engine/engine.go).

N4: run on the real kernels. Round 2's rig, reused: the same Ubuntu-archive .debs, whose hashes tools/qemu_inner_r4.sh compares with round 2's verifier before booting (pr2-r2/verify/v2/logs/kernel-debs.sha256) and which it refuses to boot if they differ — they are identical. qemu-system-aarch64 (TCG), qinit as PID 1, and this head's ./engine/iouring test binary built with -race. Source: qemu/VERDICT.txt, from tools/qemu_verdict_r4.py.

kernel the probe's answer the kernel executed the forced reap TestReapOnTheRunningKernel/probe_matches_the_kernel
Ubuntu 5.15.0-191 rejected: cqe.res=-22 (EINVAL); the kernel predates Linux 5.19 no PASS, on its rejecting branch, which had never run anywhere: rejecting branch: the kernel failed the forced reap (TransplantReapFailed=1), and the conn left after its held response (TransplantHeld=1, adopted 1)
Ubuntu 5.19.0-50 (control) accepted yes PASS
Ubuntu 6.8.0-138 (control) accepted yes PASS
  • On each kernel the five probe tests PASS (TestAsyncCancelProbeClassifies, TestAsyncCancelProbeOnThisKernel, TestWorkersCarryTheAsyncCancelProbe, TestReapOnTheRunningKernel, TestNoReapWithoutAsyncCancelFlags), round 4's own subtests PASS, every engine is at one io_uring worker, and there is no WARNING: DATA RACE.
  • What the run found. TestHandoffHasNothingInFlight/async failed on 5.15 — "0 of 128 busy conns were handed off, want all" — and it fails there on 272bcba too (qemu/diag/, arm base-b: its async subtest FAILs, this head's PASSes). Without the cancel flags a promoted async conn cannot be reaped, so it is never offered for the hand-off and stays on io_uring (asyncTransplantEligible, R1); the subtest wanted every conn handed off on every kernel. 5ce4986 makes it want none where the probe did not find the flags, with the loss checks unchanged, and adds async_without_cancel_flags, which takes the probe's answer as a rejection so that branch runs on every kernel, CI's included.
  • Two failures under QEMU are the apparatus, and 272bcba has them too. The follow-up that shows it (qemu/diag/, tools/qemu_build_diag_r4.sh + qemu_run_diag_r4.sh) runs the same two halves for 272bcba's -race binary, this head without -race, and this head with a diagnostic 10 s client read deadline (never committed):
    • TestReapedRecvLeavesNoLinkOrBuffer/provided_buffer_goes_back on 5.15 and 6.8: register pbuf_ring: io_uring_register pbuf_ring: invalid argument. 272bcba's binary fails the same subtest on the same two kernels (base-a).
    • The engine-level load tests' read_timeout errors on 5.15 and 5.19: TCG emulates the machine in software. Without -race all four pass on 5.15 (nr-b), and with the 10 s deadline they pass on 5.15 and 5.19 (dl10-b); 272bcba fails them there too. On 6.8 all four pass as they are.

N5: done — I stopped naming it. The mutation set no longer claims that test as a killer anywhere: not in MUTANTS.tsv, not in mut/VERDICT.txt, not in the PR body.

  • m04's must-fail tests are now TestTransplantNeverHandsOffArmedRecv (synthetic completions) and TestHandoffHasNothingInFlight/async (every connection is promoted, and a promoted connection always leaves through the reap: W2 = 128 in both m04 runs). tools/analyze_mut_r4.py now accepts a subtest name in a must-fail entry.
  • The same reasoning applies to two more round-1 mutants that named the engine-level load tests for this branch, so I re-scoped them too and said so in mut/VERDICT-NOTE.txt: m01 (the same gate, removed instead of bypassed) names the same two killers; m23 (the same gate with the witnesses blinded) names the three tests whose checks read the hand-off itself — TestTransplantNeverHandsOffArmedRecv and TestTransplantHandoffInFlightCounts{TryTransplant,FinishAsyncTransplant} ("the R0 gate is gone") — so it is caught with no client involved. TestStaleRecvDataCounted still runs in all three mutants' -run sets and the verdict reports what it did; it is just not required.
  • The census behind it, over every run of rounds 2 to 4 that ran the test: 18 runs, exposure 0 to 88 of 128 connections, 0 in 6 of them (mut/EXPOSURE.txt, tools/exposure_census_r4.py).
  • In this round's run all three are killed by their new killers: m01 and m04 by TestTransplantNeverHandsOffArmedRecv ("handed off 1 time(s) with the linked RECV armed") and TestHandoffHasNothingInFlight/async ("TransplantHandoffInFlight read 40 at a hand-off, want 0 at every one"), and m23 by the three hand-off tests. TestStaleRecvDataCounted happened to FAIL under all three this time; the verdict reports that under other-non-PASS and does not rely on it.

N6: done. The PR body now opens with: "Squash-merge this PR. Its branch carries 683c7bb, "DO NOT MERGE: plant the no-HOLD/REAP engine …", the deliberate red commit round 3's R6 asked for, reverted by 272bcba. A squash keeps that commit out of main's history; a merge commit or a rebase would carry it in (N6)."

Checks on this head.

  • Failing first (ff/VERDICT.txt): quoted under N1-N3 above; the seam-only control passes.
  • ci.yml is unchanged (git diff 272bcba HEAD -- .github/ is empty). The three GOOD arms were re-run on this head from the same verbatim run: blocks, one container each (interlock/5ce4986/VERDICT.txt, tools/il_verdict_r4.py): the witness step passes 29/29 with SKIP 0 and 8 engines at workers=1 (one more than round 3 — async_without_cancel_flags builds one); the one-worker leg passes 2/2 with two engines at workers=1 and both summaries at err=0; the race step passes with T4 1/1 and cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=2109529 err=0. The other five arms are round 3's and stand.
  • Mutation (mut/VERDICT.txt, tools/analyze_mut_r4.py): 85 mutants (round 3's 74 re-derived on this head plus q01-q11 for round 4's behaviour), one container each. 85 killed, 0 survivors, each by the tests named for it and for the registered reason where one is. The unmutated controls passed 29/29 with 44 subtests, 2/2 and 1/1; every io_uring run printed workers=1 for all 8 engines and every T4 RESULT line read wE=2 wI=2.
  • Suites, this head against 272bcba (suites/COMPARE.txt, tools/compare_suites_r4.py), -race -v: ./engine/iouring at 8 MiB (154 PASS, 5 SKIP) and at 128 MiB with CELERIS_REQUIRE_IOURING_WORKERS=1 (159 PASS), ./adaptive with the CI -skip at 8 MiB (72 PASS, 4 SKIP) and at 128 MiB with CELERIS_REQUIRE_UPSWITCH=1 (76 PASS) — the same tallies on both trees. 0 items PASS on 272bcba and fail to PASS here; no FAIL, no WARNING: DATA RACE, no panic in any of the 8 runs. The only differences are round 4's six new subtests, absent on 272bcba and PASS here in both memlock cells.
  • Lint as the lint job runs it is clean, each linter with its planted control (lint/, tools/run_lint.sh).
  • CI on this head is green: runs 35474330959 (CI) and 35474327611 (analysis), all 11 jobs at the first attempt, none re-run (ci/5ce4986/SUMMARY.txt, tools/summarize_ci_r4.py, over the real job logs in ci/5ce4986/logs). Witness step: want 29, ran 29, passed 29, SKIP lines 0, engines 8 at workers=1: 8, with 44 subtests PASS at memlock 8192 KiB. T4: want 1, ran 1, passed 1, SKIP lines 0, RESULT lines 1, at two workers with err=0: 1, RESULT cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=616472 err=0 w1=0 w2=0 doubleclaim=0 holdrescued=0 reapfailed=0, all six switches moving and adopting 256. One-worker leg: want 2, ran 2, passed 2, SKIP lines 0, io_uring engines 2 at workers=1: 2, summaries 2 with err=0: 2. Adaptive package 76 PASS, 0 FAIL, 0 SKIP. On the runner's kernel (6.17.0-1022-azure) the probe answered accepted and the forced reap ran, and async_without_cancel_flags PASSed there with 0 of 128 promoted conns handed off.

…e async hand-off subtest demands (celeris#681 M1)

TestHandoffHasNothingInFlight/async chose between "all 128 promoted async
conns handed off" and "none of them" from e.asyncCancelFlags alone. That
boolean is false for every probe answer except an acceptance, so a probe
that got no answer -- its private ring failing with EMFILE or ENOMEM, which
celeris#681 N1 made a per-New event -- silently downgraded the subtest's
strongest assertion into one any run with the reap off passes trivially.
The subtest is the registered deterministic killer of mutants m01 and m04,
so the mutation claim inherited the weakness.

The branch is now chosen by the running kernel: from 5.19, which is where
IORING_ASYNC_CANCEL flags exist, the strong branch is demanded and an
engine whose reap is off fails the subtest, naming the probe's class from
New's own "io_uring engine selected" record rather than probing a second
time. Only a kernel that predates 5.19 may take the weak branch.
async_without_cancel_flags, which forces the no-flags path on every kernel,
is unchanged.
…, and why the two probes in TestReapOnTheRunningKernel stay (celeris#681 N-a, N-d)

runAsyncCancelProbe, asyncCancelProbeSubmit, asyncCancelProbeWait and
asyncCancelProbeTimeout are read inside probeAsyncCancelFlagsCached's
asyncCancelMu critical section and replaced by tests that hold no lock.
Nothing in engine/iouring calls t.Parallel(), so nothing races today; the
constraint is now written where the seams are declared, with the grep that
checks it, rather than enforced with a mutex-taking setter per seam.

TestReapOnTheRunningKernel's two probeAsyncCancelFlagsCached calls can be
two probes since N1 cached only an answer. They stay: neither subtest
weakens a check on the answer it got, and probe_matches_the_kernel Fatalfs
when the answer disagrees with what the kernel does with a real reap.
EngineMetrics.TransplantReapUnsupported had a 29-character fragment line
mid-paragraph; asyncTransplantEligible's comment had a 131-character line
among neighbours that wrap at about 78. Text unchanged.
@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Round 5 is pushed at edca5d2. It answers M1 and N-a to N-e. The PR body is updated (the round-5 paragraph and residual 7). Every figure here comes from a verdict file named beside it, each regenerated by a saved script; round 5's evidence is under evidence/celeris-657/impl/pr2-r5 (manifest SHA256SUMS, index INDEX.txt, base directory stated in both).

Commits: 8e09151 (M1), 26f9cb9 (N-a, N-d), edca5d2 (N-b).


M1 (minor): fixed. The kernel decides what the subtest demands, not the probe's answer.

The rule I implemented, in TestHandoffHasNothingInFlight:

running kernel probe's class what /async demands
5.19 or later accepted the strong branch: all 128 promoted async conns handed off
5.19 or later rejected, no answer, or unexpected FAIL at the gate, naming the class — a kernel whose version has the flags must have the reap on
before 5.19 any the weak branch if the reap is off (e.asyncCancelFlags), the strong one if it is on
any kernel async_without_cancel_flags (forced rejection) the weak branch — unchanged

kernelHasFlags is e.profile.KernelMajor > 5 || (major == 5 && minor >= 19), read from the engine's own probed profile — the same numbers New gives logAsyncCancelProbe. The three non-accepted classes are not told apart in the decision (at 5.19+ each one means the reap is off where it should work), but the failure message names them, and it names them from New's own records rather than by probing again: the subtest gives the engine a JSON logger and reads async_cancel_probe out of the io_uring engine selected record plus the reason from the async cancel flags … record. That was deliberate — a second probe need not give the answer this engine was built with, which is exactly N1's point.

Failing first — one seam, engine/iouring/probe.go's runAsyncCancelProbe returning asyncCancelNoAnswer and nothing else (ff/ovl/SEAM-HEAD/probe.go.seam.diff, generated by tools/mk_ff_r5.py), four arms, one container each, on this 7.0.12-linuxkit kernel, all running the same -run '^TestHandoffHasNothingInFlight$' (ff/VERDICT.txt, tools/ff_verdict_r5.py):

arm tree /async what it shows
FF-BASE-PLAIN 5ce4986, no seam PASS the strong branch: all 128 handed off (probe accepted)
FF-BASE-SEAM 5ce4986 + the seam PASS trivially — the weak branch ran, 0 of 128 promoted async conns handed off. The defect: a vacuous check on a 7.0 kernel
FF-HEAD-SEAM this head + the same seam FAIL at the M1 gate, naming the class no answer and the seam's reason
FF-HEAD-PLAIN this head, no seam PASS the strong branch again: all 128 handed off (probe accepted, reap=true)

The two base arms are the discriminating pair: the same tree, the same test, the seam the only difference, and it turns a real check into a vacuous one. The head's failure message is

fd_lifetime_engine_test.go:374: the engine's cancel-flags probe did not find the flags accepted on kernel 7.0.12-linuxkit (7.0), which has had them since 5.19, so the hand-off's reap is off and no promoted async conn is offered for it: WARN "async cancel flags probe got no answer from the kernel: the io_uring→epoll hand-off will not cancel an armed recv (celeris#657)" reason="celeris681 FF forced: the probe got no answer"; New's probe class="no answer". […]

On real kernels. Round 2's QEMU rig again — the same Ubuntu-archive .debs, whose hashes tools/qemu_inner_r5.sh compares with round 2's verifier before booting (they are IDENTICAL), qemu-system-aarch64 (TCG), qinit as PID 1 — with two -race binaries this time: the head, and the head built with the same forced-no-answer overlay (qemu/VERDICT.txt, tools/qemu_verdict_r5.py). Both run only the subtest M1 changed.

kernel head.test /async seam.test /async (probe forced to no answer)
Ubuntu 5.15.0-191-generic (a version with the flags: false) PASS PASS — the weak branch is legitimate here
Ubuntu 5.19.0-50-generic (a version with the flags: true) PASS FAIL — at the M1 gate
Ubuntu 6.8.0-138-generic (a version with the flags: true) PASS FAIL — at the M1 gate

That is the rule end to end: the weak branch survives exactly where the kernel genuinely predates 5.19, and a probe with no answer is fatal from 5.19 on. (/sync fails on 5.15 and 5.19 under TCG with read_timeout client errors and no witness movement — the apparatus limit round 4 diagnosed in pr2-r4/qemu/diag; on 6.8 it passes. It is reported and not gated.)

The mutation claim it carries. The subset re-run on this head is every mutant whose must-fail list names TestHandoffHasNothingInFlight — 7 of them, derived from the TSV by tools/mut_subset_r5.py, which refuses unless m01 and m04 are in it: m01_no_r0_gate, m03_no_r0_gate_async_site, m04_reap_hands_off_anyway, r02_park_check_inverted, r12_doubleclaim_planted, r14_reapfailed_planted, r16_holdrescued_planted. No mutant overlays a _test.go file, so no other mutant's tests differ between the two heads. 7 of 7 killed, survivors: none, each by the tests registered for it. m01 and m04 are still killed by TestHandoffHasNothingInFlight/async with the strong branch running, not the weak one: “fd_lifetime_engine_test.go:400: TransplantHandoffInFlight read 40 at a hand-off, want 0 at every one: a conn left with an op still able to resolve its fd”.


N-a (nit): documented, not guarded — and here is why. The four seams (runAsyncCancelProbe, asyncCancelProbeSubmit, asyncCancelProbeWait, asyncCancelProbeTimeout) now carry the constraint where they are declared: they are read inside probeAsyncCancelFlagsCached's asyncCancelMu critical section and written by tests that hold no lock, so a test may replace one only from its own goroutine, with no other goroutine able to reach a probe — never from a t.Parallel() test, never while an engine it does not own is being built. The comment names the check: git grep -n 't\.Parallel()' -- 'engine/iouring/*_test.go' is empty (positive control: the same grep finds engine/provider_test.go). I chose documenting over guarding because guarding means a mutex-taking setter and a save/restore helper for each of the four, in production code, for a constraint one grep checks and nothing in the package violates. If a parallel test ever lands in this package, the grep is the thing that has to be re-run.

N-b (nit): fixed. engine/engine.go's TransplantReapUnsupported paragraph is reflowed as a whole (the 29-character fragment is gone; longest line now 74), and transplant_source.go's asyncTransplantEligible comment is reflowed (the 131-character line is gone; longest line now 78, against neighbours at 77–78). Text unchanged in both — docs: commit edca5d2.

N-c (nit): added. Residual 7 in the PR body now states round 4's per-New cost: a kernel that never answers, or answers with something the probe does not recognise, is re-probed at every iouring.New, serialised on a process-global mutex, with a log record each time (WARN from 5.19, INFO below), each probe opening and closing a private ring, and in the remote no-CQE case holding that mutex for up to 500 ms (a wait to a 500 ms deadline, repeated once if a signal cuts it short). It also says what the cost buys — cached, one EMFILE or ENOMEM kept the reap off for the life of the process — and that no kernel in this PR's evidence takes the path.

N-d (nit): I leave it, and said so in the code. TestReapOnTheRunningKernel's two probeAsyncCancelFlagsCached calls can now be two probes. Neither subtest downgrades a check on the answer it got: probe_answer runs the branch its answer names, and probe_matches_the_kernel compares its answer with what the kernel does with a real reap and t.Fatalfs on any mismatch — so two probes that disagree fail loudly. The test's doc comment now says this, and points at M1 as the one place where an answer did choose a weaker check.

N-e (nit, pre-existing): recorded as accepted. ci.yml is untouched. git diff 5ce4986..HEAD -- .github/workflows/ci.yml is empty and both blobs hash to 34b53f1fca16, so only the three GOOD interlock arms were re-run. Why accepted rather than fixed:

  • The gap is narrower than "subtests are not counted". The step's skipped term greps ^[[:space:]]*--- SKIP, which matches an indented line, so a skipped subtest already fails the step — measured, round 2's W-SUBSKIP arm: want 29 ran 29 passed 29 SKIP 1 → step rc 1 (pr2-r2/fix/interlock/48dfe1a/VERDICT.txt). A failing subtest fails its parent, so passed drops below want. What is uncovered is a subtest deleted or renamed with its parent still passing.
  • That case is covered off the runner, every round, by the suite comparison, which counts subtests by name: this round ./engine/iouring at 8 MiB 115 subtest PASSes and at 128 MiB 115, identical on both trees — a deleted subtest shows up there as a base PASS that is ABSENT on the head, and the tool flags it.
  • Closing it on the runner means changing ci.yml, and the standing rule for that is the full interlock both ways: all 28 arms, not the 3 GOOD ones. In the last round before merge I would rather not spend the change budget there, and I would rather not land a ci.yml edit with a partial proof. Filing it is the honest option if you want it closed; say the word and I will open it against PR 3, which touches this step anyway.

Checks on this head (all local runs: one container at a time, golang:1.27, --cpus 4, seccomp unconfined, -race -v, count=1; conc_max=1 recorded in every log).

  • Failing first (ff/VERDICT.txt): the table above; FAILING_FIRST PASS.
  • QEMU, three real kernels (qemu/VERDICT.txt): the table above; QEMU_M1_RULE PASS, 0 WARNING: DATA RACE lines on all three.
  • Mutation (mut/VERDICT.txt, tools/analyze_mut_r5.py, subset derived by tools/mut_subset_r5.py): 7 of 7 killed, survivors: none; one container each. The unmutated control passed 29 of its 29 top-level tests with 44 subtests, none non-PASS, every engine at workers=1.
  • Suites, this head against 5ce4986 (suites/COMPARE.txt, tools/compare_suites_r5.py): ./engine/iouring at 8 MiB (154 top-level PASS, 5 SKIP, 115 subtest PASS) and at 128 MiB with CELERIS_REQUIRE_IOURING_WORKERS=1 (159 PASS, 0 SKIP, 115 subtest PASS) — byte-for-byte the same tallies on both trees, 0 items differing in either cell, 0 PASS->non-PASS, no FAIL, no WARNING: DATA RACE, no panic in any of the 4 runs. Round 5 touches no adaptive code or test, so round 4's adaptive suite cells are not repeated; the adaptive job is read from CI below.
  • Interlock, the three GOOD arms on this head (interlock/edca5d2/VERDICT.txt, tools/il_verdict_r5.py), steps extracted verbatim from ci.yml at this head with the same run_sha256 values round 4 recorded: the witness step want 29 ran 29 passed 29 SKIP 0 engines 8 at1 8; the one-worker leg want 2 ran 2 passed 2 SKIP 0 engines 2 at1 2 summaries 2 err0 2; the race step want 1 ran 1 passed 1 SKIP 0 RESULT 1 clean 1 with T4 cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=2134271 err=0. All three expected pass got pass.
  • Lint as the lint job runs it, each linter with its planted control (lint/, tools/run_lint.sh): clean — 10 golangci-lint cells (5 modules x linux/amd64+arm64) at issue_lines=0, mage-compile and mage CheckRelease at rc=0, actionlint 0 lines, zizmor rc=0; and all three planted controls fired: CONTROL planted unquoted variable rc=1 sc2086=1; CONTROL planted-defect rc=1 errcheck=1 ineffassign=1 unused=1; CONTROL zizmor planted rc=13 artipacked=2.
  • CI on this head: green. Runs 35476784679 (CI), 35476782587 (PR fix(iouring): hand a connection to epoll only when no read can still resolve its fd (celeris#657) #681), every job at attempt 1, none re-run (ci/edca5d2/SUMMARY.txt, tools/summarize_ci_r5.py, over the REAL job logs in ci/edca5d2/logs). Witness step: want 29, ran 29, passed 29, SKIP lines 0, engines 8 at workers=1: 8 — and its subtest count, which the step itself does not check (N-e): 29 top-level PASS and 44 subtests PASS, nothing non-PASS. T4: want 1, ran 1, passed 1, SKIP lines 0, RESULT lines 1, at two workers with err=0: 1, RESULT cycles=3 conns=256 wE=2 wI=2 conns_per_ring=128 ok=1030659 err=0 w1=0 w2=0 doubleclaim=0 holdrescued=0 reapfailed=0 held=767 reaps=59. One-worker leg: want 2, ran 2, passed 2, SKIP lines 0, io_uring engines 2 at workers=1: 2, summaries 2 with err=0: 2. Adaptive package: 76 top-level PASS, 16 subtests PASS, nothing non-PASS. 0 WARNING: DATA RACE lines across every job. On the runner's kernel the strong branch is the one that ran: kernel=6.17.0-1022-azure (6.17, a version with the flags: true) forced_no_flags=false reap=true; New's probe class="accepted".

@FumingPower3925
FumingPower3925 merged commit 4770d07 into main Sep 20, 2026
12 checks passed
@FumingPower3925
FumingPower3925 deleted the fix/celeris-657-fd-lifetime branch September 20, 2026 00:01
FumingPower3925 added a commit that referenced this pull request Sep 26, 2026
…ines (celeris#662, celeris#657)

The pause keeps a listener open for about 1.5 s with TCP_DEFER_ACCEPT
cleared, so for that long the engine a switch is leaving keeps its
SO_REUSEPORT share of NEW connections. Nothing pinned where those go, and
that is exactly what failed this PR's base-vs-branch suite gate on
2026-09-18: with the linger in and celeris#657 open, an async 2048-connection
ramp left 972-1170 conns on the outgoing epoll after 20 s, err=0. Not a loss
-- placement. celeris#657 is now fixed (#681, #687) and this is the test that
holds the joint property, so the interaction cannot regress silently.

TestLingerArrivalsReachTheIncomingEngine, in engine/epoll and engine/iouring,
with an async handler and the default configuration:

  * BeginPauseAccept, then wait for the socket witness to read the option off
    every listener -- the linger has begun;
  * PHASE A: six connections arrive, are served, and go silent while NO drain
    is set. Their dispatch goroutines park with nothing owed for them (epoll:
    askAtPark returns early with no drain; io_uring: runAsyncHandler's claim
    is not reached). This is the flap case, a second switch inside the first
    one's linger;
  * StartTransplant, still inside the linger;
  * PHASE B: six more arrive, are served, and go silent with the drain set.
    They join the live set at the TAIL, past the sweep cursor -- the case the
    cycle rule exists for (celeris#657 R2 MAJOR-1);
  * then every listener closes and the engine must hold none of them.

It asserts what a client and an operator can each check: no dial refused and
every arrival accepted by THIS engine (OnConnect is the witness), every
arrival answered 200, every one adopted by the target, ActiveConnections back
to 0, and the hand-off ledger balanced (io_uring also requires
TransplantHandoffInFlight 0).

The target serves what it adopts. A target that only holds the descriptor
cannot tell a connection that was MOVED from one that was DROPPED, because
either way its client never gets an answer -- the first draft of this test
read 12 timeouts and 12 adoptions and could not say which had happened.

Both controls lose, measured in one container each, mutation restored and
sha256-verified (evidence celeris-662/r6/rebase/logs/c05..c08):

  * linger 0 (celeris#662 off): 12 of 12 dials refused, 0 accepted, 0 served,
    on both engines.
  * the sweep and the park-boundary ask removed (celeris#657 off): epoll
    adopts 0 of 12 and still holds 12; io_uring adopts 6 of 12 and still
    holds the 6 of phase A. io_uring's phase B moves without the sweep
    because its own park claims it, which is why phase A is in the test: it
    is the only shape on that engine that the sweep alone carries.

CI: the name joins the celeris#662 linger step's tally, which runs both
packages with -v, forbids every skip and asserts the exact PASS count.
FumingPower3925 added a commit that referenced this pull request Sep 26, 2026
…on the two ramp tests (celeris#662)

The quarantine comment was written when celeris#657 was open and said the
two Ramp tests fail on it. Both halves of #657 are on main now (#681, #687),
so that paragraph is stale in the tree this branch produces.

Measured on the rebased head, 3 containers, golang:1.27, --cpus 4, memlock
128 MiB (4 io_uring workers), seccomp unconfined, -race -count=1, one package
per container (evidence celeris-662/r6/rebase/logs/s06,s07-rep1,s08-rep2,
regenerated by tools/tally.sh):

  TestRampH1Sync  PASS 3/3
  TestRampH1Async PASS 3/3
  every high phase, both variants: epoll=0 of 2048, settled_in 2.50-2.55 s
  err=0 over 6,755,764 requests; err census total=0, no class at all

That is the registered prediction on celeris#687 (comment 5749130618) held:
the residue that failed this PR's own G5 gate -- 972-1170 of 2048 left on
the outgoing epoll after 20 s -- is what #687's sweep removes.

The names stay on -skip regardless. The lift rule is >= 6 GitHub-hosted
re-runs of identical bytes, the same rule #687 applied to
TestBidirectionalFlapAsync, and a 2048-connection ramp on a shared runner is
the shape celeris#686's dial-time resets were seen in. Lift it from a green
PR run.
FumingPower3925 added a commit that referenced this pull request Sep 28, 2026
… resolve it: close paths, hijack, shutdown (celeris#685)

Fixed files are off, so a recv or send SQE names its descriptor by number
and the kernel resolves the number when it issues the op. finishClose and
finishCloseDetached queued the ops' cancels and closed the descriptor at
once, while an op could still be issued: a recv still in the SQ ring (a
promoted connection's re-arm), or one linked behind a SEND that the kernel
issues only after the SEND completes. A sibling worker given the freed
number before this worker's next enter then had its new connection's
request read by the old recv, dropped as stale_recv_data_closed, and the
client was never answered (celeris#715: 80 of 80 in #781's trial).

The rule, as #681 applied it to the hand-off: a close path does not release
the number while the kernel owes an op on it (fdOwed: kernelInflight > 0,
which counts every recv and send from the moment its SQE is written). It
does everything else as before and hands the descriptor to its
pendingRelease entry; drainPendingRelease closes it where it releases the
connState, at the last owed op's terminal CQE. The socket's read side is
shut down too (SHUT_RD on the H1 fast path, SHUT_RDWR where the path
already half-closed), so an owed recv ends as soon as it is issued, even a
linked one no cancel can find, and even where cancels fail (#682). No
io_uring_enter is added; the close(2) moves one iteration later. The worker
does not park while such a close is outstanding.

hijackConn keeps the socket open under the hijacker, so it submits its
cancels before handing the socket over when an op is owed (a multishot recv
stays armed across its request). Worker shutdown ends the ops owed on
connection descriptors (cancel, SHUT_RDWR, run the ring until their
terminal CQEs, bounded at 250 ms) before it closes any, and before
shutdownDrivers, which relies on nothing being submitted after it.

EngineMetrics gains CloseFDDeferred (a rate) and CloseFDForced (the backstop
closing a descriptor with an op still owed; must stay 0). The recv-theft CI
job gets a detector control: with fdOwed forced false the three judging
trials must fail every run.

Fixes #685
Refs #715
FumingPower3925 added a commit that referenced this pull request Sep 28, 2026
… resolve it: close paths, hijack, shutdown (celeris#685)

Fixed files are off, so a recv or send SQE names its descriptor by number
and the kernel resolves the number when it issues the op. finishClose and
finishCloseDetached queued the ops' cancels and closed the descriptor at
once, while an op could still be issued: a recv still in the SQ ring (a
promoted connection's re-arm), or one linked behind a SEND that the kernel
issues only after the SEND completes. A sibling worker given the freed
number before this worker's next enter then had its new connection's
request read by the old recv, dropped as stale_recv_data_closed, and the
client was never answered (celeris#715: 80 of 80 in #781's trial).

The rule, as #681 applied it to the hand-off: a close path does not release
the number while the kernel owes an op on it (fdOwed: kernelInflight > 0,
which counts every recv and send from the moment its SQE is written). It
does everything else as before and hands the descriptor to its
pendingRelease entry; drainPendingRelease closes it where it releases the
connState, at the last owed op's terminal CQE. The socket's read side is
shut down too (SHUT_RD on the H1 fast path, SHUT_RDWR where the path
already half-closed), so an owed recv ends as soon as it is issued, even a
linked one no cancel can find, and even where cancels fail (#682). No
io_uring_enter is added; the close(2) moves one iteration later. The worker
does not park while such a close is outstanding.

hijackConn keeps the socket open under the hijacker, so it submits its
cancels before handing the socket over when an op is owed (a multishot recv
stays armed across its request). Worker shutdown ends the ops owed on
connection descriptors (cancel, SHUT_RDWR, run the ring until their
terminal CQEs, bounded at 250 ms) before it closes any, and before
shutdownDrivers, which relies on nothing being submitted after it.

EngineMetrics gains CloseFDDeferred (a rate) and CloseFDForced (the backstop
closing a descriptor with an op still owed; must stay 0). The recv-theft CI
job gets a detector control: with fdOwed forced false the three judging
trials must fail every run.

Fixes #685
Refs #715
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working engine/iouring io_uring engine specifics

Projects

None yet

1 participant