Skip to content

feat(engine): count the requests an io_uring hand-off drops, and the hand-offs made with an op in flight (celeris#657) - #676

Merged
FumingPower3925 merged 14 commits into
mainfrom
fix/celeris-657-instruments
Sep 19, 2026
Merged

FumingPower3925 merged 14 commits into
mainfrom
fix/celeris-657-instruments

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This is PR 1 of 3 for #657. It adds instruments only and changes no behaviour. It makes the request loss diagnosed in #657 (diagnosis comment) countable. The fix that follows in PR 2 can then be judged by these counters reading zero, and this PR first shows they are non-zero on today's engine.

Refs #657

Why these counters exist

When io_uring hands a sync connection to epoll (a revert), tryTransplant dups the fd, closes the original and hands the dup over. The connection's next RECV can still be issued or completed after that. When it reads, it takes a request, and through a recycled fd number that request can belong to a different connection. Its CQE still carries the old (fd, generation), so staleConnCQE drops it as stale, and nothing counted it. The #624 hand-off ledger (TransplantDetached == TransplantAdopted) keeps balancing while requests go missing. The diagnosis comment has the measurements and the mechanism.

The third instrument is the ramp test's error census. As the correction comment explains, recordErr capped the whole tally on map size, and its keys carry the client's ephemeral port, so it went quiet after about 200 connections. The same census is how anyone would measure the fix.

What this adds

EngineMetrics field counted at meaning
StaleRecvDataTransplanted staleConnCQE: a stale recv CQE with res > 0 whose (fd, gen) this worker handed off a request read by a recv that outlived its hand-off. This is the #657 loss.
StaleRecvDataClosed same, for an identity this worker closed or hijacked usually a client's bytes that raced a server-side close. Not always benign: after a Hijack the socket lives on under the hijacker's net.Conn, and a recv that reaches the kernel after a close resolves the fd number, which a new connection may hold by then. Either way a live client's request can land here.
StaleRecvDataUnattributed same, with nothing registered for the identity a gap in attribution
TransplantHandoffInFlight tryTransplant and finishAsyncTransplant, once the hand-off is committed hand-offs made while recvArmed, kernelInflight != 0 or zcNotifPending. This is the precondition of every loss above, and PR 2 must take it to 0.
  • The class is read before noteStaleTerminalOp retires the closedOps entry. Counting after the retire would turn every identity into Unattributed, and a mutant pins this ordering.
  • The three StaleRecvData counters are not exhaustive. A stale completion is recognised only because its (fd, generation) differs from the fd's current occupant. One whose pair matches the occupant, a generation collision, takes staleConnCQE's live branch and none of them counts it. Generations come from one process-wide 32-bit sequence, so this needs the sequence to wrap while the op is in flight. The residual predates this PR, and the field doc now says so.
  • Both hand-off sites register the identity through noteHandedOffInflight, which is noteClosedInflight plus a handoff flag on the closedOpsEntry. Kernel accounting and the release gate are unchanged. The flag sits in inflight's padding, so on 64-bit platforms the entry stays 32 bytes. This is pinned by TestClosedOpsEntryStaysThirtyTwoBytes and measured at compile time for linux/amd64 and linux/arm64. On 32-bit platforms there is no padding there, and the entry grows from 16 to 20 bytes (measured for linux/386 and linux/arm).
  • The in-flight predicate keeps all three terms. With consistent accounting kernelInflight != 0 implies the other two, but a generation-collision misroute can take kernelInflight to 0 while a recv is still armed or a SEND_ZC notification is still pending. The three terms are also the condition PR 2 refuses a hand-off on. The unit test checks each term on its own.
  • A multishot F_MORE data CQE counts once per CQE. A unit test pins this. No integration run has produced one yet: 0 of 1,024 stale recv completions in the join proof, and 0 in the T4-shape campaign.
  • All four fields are cumulative and io_uring-only, and read zero on epoll and std. adaptive.Engine.Metrics() sums both sub-engines, because by the time the stale CQEs arrive the engine that made the hand-off is the standby.
  • They are direct atomic adds on the stale-CQE and hand-off paths only. Nothing runs on the per-request path.
  • Census (adaptive/ramp_transplant_test.go): recordErr now counts every error, in total and per port-free class. Only the verbatim exemplars are capped at 200, and a key that is already held keeps counting. dumpErrSamples prints the class tally first.

There are 13 new tests, one per counter or path, with no subtests:

  • engine/iouring/handoff_loss_test.go has 8. Two of them drive tryTransplant and finishAsyncTransplant on a ring-backed worker, then feed that identity's recv completion. One is a negative control: completions that read nothing, stale sends and a live connection's own data must not count. One feeds a multishot recv's F_MORE completions.
  • engine/iouring/handoff_loss_metrics_test.go has 2.
  • engine/iouring/closed_ops_entry_size_test.go has 1 (linux/amd64 and linux/arm64 only).
  • adaptive/handoff_loss_metrics_test.go has 1.
  • adaptive/ramp_err_census_test.go has 1.

The existing TestMetricsCarriesEveryFieldReflectively covers the new fields.

CI interlock. The unit job runs ./engine/iouring without -v, so a skip there prints nothing, and five of the eleven engine/iouring tests can skip. A new step in the unit job re-runs all eleven by name. It uses -v and CELERIS_REQUIRE_IOURING_WORKERS=1, which newTestRing and the worker test now honour, so an environment skip becomes a failure. It then checks an exact tally: === RUN Name and --- PASS: Name ( must each count 11, and there must be no --- SKIP line. The step's block was run verbatim, as GitHub runs shell: bash, against the committed tree and against copies with one test source edited:

arm step result
committed tree exit 0: want 11, ran 11, passed 11, SKIP 0
one test deleted exit 1: ran 10, passed 10
one test renamed exit 1: ran 10, passed 10
one test calling t.Skip exit 1: passed 10, 1 SKIP line
committed tree, io_uring refused by seccomp exit 1: the 5 ring-backed tests FAIL ("forbids skipping")

How it was checked

Base is main 985a386, and this branch is 1aca300 (985a386 plus 13 commits).

  • Every campaign ran in Docker: golang:1.27, --cpus 4, one container per observation and one at a time, with -race -v.
  • Only top-level --- PASS/FAIL lines are tallied, and workers= is checked in every run.
  • Predictions were hashed before each campaign's first run.
  • Every figure below is regenerated from the saved run logs by a named analyzer: analyze_j.py, analyze_t.py, analyze_v.py, analyze_v2.py, analyze_s.py and analyze_mut.py.
  • The evidence tree they live in is local and hashed, so the figures are quoted here rather than linked.

1. Join proof: each counter equals the loss it claims to count

One worker, sync site (campaign J). The cells are TestReverseTransplant and TestBidirectionalFlap at memlock 8 MiB, which caps io_uring at one worker; workers=1 held in 32/32 runs. There are two arms, n = 8 per test per arm, 32 runs in all:

  • JB is this branch's engine byte for byte, plus a test-file overlay that splits client errors into dial, write and read and prints the full EngineMetrics after the load.
  • JK is JB plus a probe that records every hand-off (fd, gen) and every stale completion, independently of the counters.
check result
StaleRecvDataTransplanted == client read errors 32/32 runs, 869 == 869 in total
StaleRecvDataUnattributed == StaleRecvDataClosed == 0 32/32
client errors other than read errors (dial, write) 0 in 32/32
TransplantHandoffInFlight == hand-offs the kernel confirmed as armed (a terminal stale completion arrived later for that identity, per the probe) 16/16 JK runs, 64 == 64
StaleRecvDataTransplanted == stale data CQEs the probe joined to a hand-off; unjoined data CQEs 16/16; 0
TransplantHandoffInFlight == TransplantDetached 64 == 64 in 32/32
StaleRecvDataTransplanted at the end of the load == 500 ms after a wake switch 32/32
adaptive sum == io_uring + epoll, and every epoll witness is 0 32/32
arm TestReverseTransplant: runs with loss, lost == counted TestBidirectionalFlap: runs with loss, lost == counted
JB 6/8, 81 == 81 7/8, 297 == 297
JK 7/8, 122 == 122 8/8, 369 == 369

Two workers, and the async site (campaign T). This used the T4 shape: Workers=2, 256 keep-alive connections (128 per ring) and three promote/revert cycles, at memlock 128 MiB. workers=2 held in 24/24 runs. The cells were the sync test and the same shape with every route async, so connections leave io_uring through finishAsyncTransplant. There were two arms (engine byte for byte, and with the probe), n = 6 per arm per cell:

check sync (12 runs) async (12 runs)
StaleRecvDataTransplanted + Unattributed == client read errors; dial = write = 0 12/12, 2,662 == 2,662 (loss in 12/12 runs) 12/12, 0 == 0 (no request lost in any run)
TransplantHandoffInFlight == probe hand-offs with an op in flight == kernel-confirmed armed; TransplantDetached == probe hand-offs 6/6 (1,927 hand-offs, all at tryTransplant) 6/6 (278 hand-offs, all at finishAsyncTransplant)
StaleRecvDataTransplanted == probe-joined stale data CQEs; unjoined 0 6/6 6/6

All registered predictions held. On this engine every hand-off is made with an op in flight: 0 of 2,205 probe hand-offs were quiet. So the integration runs cannot tell "count the armed hand-offs" from "count every hand-off". The unit tests' quiet arms at both sites, with mutant m21, carry that distinction until PR 2 makes quiet hand-offs the norm. The async site has not yet shown a lost request at integration level; its counter path is the sync site's (noteHandedOffInflight), and PR 2's failing-first campaign, on main plus this PR, is where it meets non-zero loss.

This is the positive control the fix needs. On an engine that behaves like main, every revert hand-off is made with its recv armed, and requests were lost in 28 of 32 one-worker runs and 12 of 12 two-worker sync runs. One instrument-check run was made before the pre-registration, on TestReverseTransplantAsync; it is not a gate cell.

2. No behaviour change

The gate was registered weaker than the design decision asked. The decision asked for "verdict rates within the A/A floor of the unchanged base". This PR registered Fisher's exact p ≥ 0.05 as the gate and kept the A/A floor only as calibration. Under the decision's criterion, campaign V failed TestBidirectionalFlap: the rate difference was 0.188, above V's largest same-bytes A/A difference, 0.125.

The pre-registered replication on identical bytes (V2) and the pooled figures:

test V, n = 16 per arm V2, n = 16 per arm pooled, n = 32 per arm
TestReverseTransplant M 15 vs B 10, p = 0.083 13 vs 13, p = 1 28 vs 23, p = 0.21
TestBidirectionalFlap M 14 vs B 11, p = 0.39 11 vs 15, p = 0.17 25 vs 26, p = 1

Counts are FAILs, base (M) against this branch (B), with Fisher's exact test, two-sided. In V2 the Flap difference reversed: the branch failed more. The pooled difference is 0.031. On identical bytes the same arm moved by up to 0.25 between V and V2 (the branch's Flap went from 11/16 to 15/16), so V's 0.125 floor, taken from 8-run halves, was too small to measure the spread. That comparison is post hoc. Read from the code, the counters add no lock, no allocation and no reordering on any path. The pooled client-error census (Mann-Whitney) gave p = 0.33 and 0.42. Both tests fail on main at one worker. That failure is #657 itself, and PR 2 is its fix.

Two registered predictions missed, both on TestBidirectionalFlap, and neither difference is significant:

  • PV6: the rate difference in V was 0.188, above the largest same-bytes A/A difference of 0.125.
  • PV2-2: in V2 the difference was 4 FAILs of 16, above the registered 3, with the branch failing more.

Full suites. The packages were ./engine/iouring, ./engine/epoll and ./adaptive, each on base and on this branch at 6d8d822. They ran at memlock 8 MiB and at 128 MiB; the 128 MiB runs set CELERIS_REQUIRE_UPSWITCH=1 and CELERIS_REQUIRE_IOURING_WORKERS=1. Each combination ran twice on identical bytes with -test.run '.', quarantined tests included, for 24 containers.

3. Mutation and lint

  • Mutation: 23 of 23 mutants were killed, each by the test meant to catch it, and the unmutated control passed 11/11 and 3/3. The mutants cover:
    • each increment, and the hand-off flag;
    • both call sites, and the count-before-retire ordering;
    • the res > 0 and op == udRecv conditions, and counting only terminal CQEs;
    • each predicate term on its own, and counting every hand-off;
    • the flag's position in closedOpsEntry;
    • the Metrics() export, the worker wiring and the adaptive sum;
    • both census regressions.
  • Lint: golangci-lint 2.13.2 with the repo's .golangci.yml, on linux/amd64 and linux/arm64, found 0 issues on the branch and 0 on the base. A defect planted in a copy of the branch is flagged (rc 1). actionlint 1.7.12 with shellcheck, and zizmor 1.30.0 offline, are clean on the workflow. A planted unquoted variable is flagged.
  • Build: go build, test-binary compile and go vet return rc 0 on linux/amd64 and linux/arm64. ./engine/iouring builds and vets on linux/386 and linux/arm.
  • gofmt: no new unformatted file.

What this PR does not do

This is PR 1 of 3:

  1. This PR: the counters and the census.
  2. PR 2, the fd-lifetime fix (face 2): no hand-off while any read can still resolve the fd. Its gate is TransplantHandoffInFlight and StaleRecvData{Transplanted,Unattributed} reading 0, with zero client errors.
  3. PR 3, placement (face 1): the connections a switch leaves on the outgoing engine.

Once this lands, the probatorium validator will report these counters in "report, don't gate" mode, so the nightly shows face 2 on both cluster architectures before the fix.

…ars (celeris#657)

recordErr guarded the whole tally on map SIZE while its keys embedded the
client's ephemeral port, so it stopped counting after ~200 connections --
repeats of a key already held included. Every figure read off it past that
point was a floor, and a "0 of kind X" reading was not concludable at all.

Counting is now unbounded and per class (errClass662 drops the addresses);
only the verbatim exemplar set is capped, and a key already held keeps
counting. dumpErrSamples prints the class census first, because it is the
only complete tally.

This is W3 of the celeris#657 instruments (PR-1). The hunk is ported
byte-for-byte from 3a195c0 on the celeris#662 branch, so rebasing that
branch onto a main that has it is a no-op for this file.

Refs #657
Refs #662
…d hand-offs made with an op in flight (celeris#657)

A reverse (io_uring->epoll) hand-off detaches a connection with its recv
still armed. Closing the original descriptor does not end an io_uring recv,
so until the cancel lands that recv keeps reading the socket epoll now owns
-- and if its SQE had not reached the kernel yet, it resolves the fd NUMBER
at submit time, which the next hand-off's dup may already have reused for a
different connection. A recv that then completes with data has consumed a
request some client is still waiting on. Its CQE carries the old (fd,
generation), so staleConnCQE drops it as stale -- with no counter at all.
The #624 hand-off ledger balanced exactly through every such loss.

Two witnesses, observation only:

* staleRecvData{Closed,Transplanted,Unattributed}: a stale recv CQE with
  res > 0, split by what closedOps holds for its identity -- a conn this
  worker closed, one it handed off, or nothing. The hand-off sites now
  register their identity through noteHandedOffInflight, which is
  noteClosedInflight plus a flag on the entry; the kernel accounting and
  the release gate are unchanged. The class is read BEFORE
  noteStaleTerminalOp, which retires the entry on the terminal CQE.

* handoffInFlight: a hand-off (tryTransplant or finishAsyncTransplant)
  that detached a conn with recvArmed, kernelInflight != 0 or a SEND_ZC
  notification pending -- the precondition of every loss above.

Neither is on the per-request path: the first runs only for stale CQEs,
the second once per hand-off. The set is nil-safe, like the #586/#591
witnesses, and is wired to Metrics() in the next commit.

Tests, one per counter and each failing if its increment is removed:
the transplanted class through tryTransplant and through
finishAsyncTransplant end to end on a real ring, the closed and the
unattributed class, a negative control (no data, a stale SEND, a live
conn's own CQE), and the in-flight witness at both hand-off sites, each
predicate term on its own.

Refs #657
…EngineMetrics

StaleRecvDataClosed, StaleRecvDataTransplanted, StaleRecvDataUnattributed
and TransplantHandoffInFlight join the #624 hand-off ledger in
EngineMetrics, wired exactly like the Transplant* counters: one engine-wide
set in the io_uring engine's metrics, a pointer to it handed to every
worker by createWorkers, read in Metrics(), and summed over both
sub-engines by the adaptive engine.

The adaptive sum matters more here than for most fields: the stale CQEs of
a revert's hand-offs arrive on the sub-engine that made them, which is the
standby by then, so reporting only the active side would read zero through
the very loss these exist to report.

io_uring-only; zero on epoll and std. This is what lets the probatorium
validator report the loss on both cluster arches before the fix lands.

Refs #657
… (celeris#657)

A unit test for W3: far more distinct error keys than errExemplarCap, as
under load, must still give an exact total and exact per-class counts, a
repeat of a retained exemplar must keep counting, and only the exemplar set
may stop growing. It fails against the old size-guarded recordErr.

Kept out of ramp_transplant_test.go so that file stays byte-identical to
3a195c0's and the celeris#662 rebase stays trivial.

Refs #657
…57 flag in its padding

The hand-off flag added for the stale-recv witness sat after the slice
header and grew the entry from 32 to 40 bytes. It fits in the padding after
inflight; an instrument should not change the layout of what it observes.

Refs #657
@FumingPower3925 FumingPower3925 added this to the v1.6.0 milestone Sep 19, 2026
@FumingPower3925 FumingPower3925 added bug Something isn't working engine/iouring io_uring engine specifics labels Sep 19, 2026
…equest, and name the stale read none of the three counts (celeris#657)

StaleRecvDataClosed was documented as benign ("nothing a live client is
waiting on"). Two routes into it contradict that. hijackConn registers
through noteClosedInflight like the close paths, and after a hijack the
socket lives on under the hijacker's net.Conn, so a recv that completes
with data before its cancel lands has read the first bytes of a
connection its client is still using. After a close, a recv that had
not reached the kernel yet resolves the fd number when it does, and a
new connection may hold that number by then.

The three counters were also documented as if they covered every stale
read. A stale completion whose (fd, generation) equals the current
occupant's, a generation collision, takes staleConnCQE's live branch
and none of them counts it. That needs the process-wide 32-bit
generation sequence to wrap while the op is in flight, and it predates
these counters. Both the EngineMetrics doc and the handoffLossStats doc
now say so.

Comments only; no code change.
…ts own (celeris#657)

TestTransplantHandoffInFlightCountsTryTransplant claimed to test each term
of `recvArmed || kernelInflight != 0 || zcNotifPending` on its own, but
its "recv armed" state also set kernelInflight to 1. A mutant that drops
`cs.recvArmed ||` survived every test.

The recv-armed state now has kernelInflight at 0. Each of the three
single-term states differs from "nothing in flight" in exactly one term.
A fifth state keeps the usual shape, a recv armed and counted.

The predicate keeps all three terms. With consistent accounting,
kernelInflight != 0 implies the other two. Every recv arm counts itself,
and a SEND_ZC keeps its count until the notification CQE that clears
zcNotifPending. A generation-collision misroute (staleConnCQE's KNOWN
RESIDUAL) can still take kernelInflight to 0 while a recv is armed or a
notification is pending. The three terms are also the condition the fix
refuses a hand-off on. noteHandoffInFlight's comment now says this.
…rows from 16 to 20 on 32-bit (celeris#657)

The handoff flag was said to keep closedOpsEntry at 32 bytes, "checked
with unsafe.Sizeof", but nothing kept the check. On 32-bit platforms the
claim was also wrong. There is no padding after inflight (an int32 plus a
12-byte slice header is 16 bytes), so the flag grows the entry from 16 to
20 bytes.

TestClosedOpsEntryStaysThirtyTwoBytes, built only for linux/amd64 and
linux/arm64, fails if the entry, or the same entry without the flag, is
not 32 bytes. The field comment now states both sizes.
…celeris#657)

The PR says a stale multishot recv counts once per data CQE, including
the intermediate F_MORE completions. No test and no integration run had
exercised that path. Every one of the 1024 stale recv completions the
join probe recorded was terminal.

TestStaleRecvDataCountsEachMultishotCompletion feeds a handed-off
identity two F_MORE data completions and then a terminal one. Each must
count as Transplanted, and the identity must stay registered until the
terminal CQE retires it.
…#657 witness tests' skips

Five of the celeris#657 witness tests can skip. Four get their ring from
newTestRing, which skipped whenever NewRing failed, with no way to forbid
it. TestWorkersShareTheHandoffLossWitnesses skipped on its own when it
could not build an engine or its workers. In a job that runs
./engine/iouring without -v, a skip prints nothing and the package still
reports ok.

newTestRing and that test now skip through skipOrFail656. It still skips
by default, and it fails when CELERIS_REQUIRE_IOURING_WORKERS=1, the
variable the `iouring` job already uses for the celeris#656 tests. The
only CI step that sets the variable today selects only the three
celeris#656 tests, so none of newTestRing's other callers change in CI.
…, skipping forbidden

The unit job runs ./engine/iouring without -v, so its output shows only
"ok" for the package. Five of the celeris#657 hand-off loss witness tests
can skip, and a skip there would pass unseen. That is how the celeris#656
leak shipped.

A new unit-job step runs all eleven witness tests by name, with the
interlock the `iouring` job established (celeris#664):
- `-v`, and an anchored `-run` built from the name list;
- CELERIS_REQUIRE_IOURING_WORKERS=1, so an environment skip is a failure;
- a tally of top-level `=== RUN   Name` lines and `--- PASS: Name (`
  lines, both of which must equal the number of names;
- zero `--- SKIP` lines at any indentation.

The PASS pattern ends with ` \(` because go test prints the elapsed time
after the name. A pattern anchored with `$` right after the name can
never match that line.

memlock stays at the runner's 8 MiB, the shape the root race step runs
these tests in. They need one ring, not a worker per 12 MiB.
… its test's comment (celeris#657)

The field doc no longer says Closed is benign: hijackConn registers
through noteClosedInflight too, and a recv that reaches the kernel after
a close can read a reused fd number. The comment on
TestStaleRecvDataCountsAClosedConn still said "the benign class". It now
says what the test checks, a close-registered identity, and points to the
field doc for the cases that are not benign.

Comment only.
…mirror struct (celeris#657)

golangci-lint's `unused` flagged the two fields of the local flag-less
struct TestClosedOpsEntryStaysThirtyTwoBytes compared against, because
they existed only for unsafe.Sizeof.

The test now reads the real struct. The entry must be 32 bytes, handoff
must sit right after inflight, and conns must stay at offset 8, where it
is without the flag. A flag placed anywhere else still fails the test.
@FumingPower3925

Copy link
Copy Markdown
Contributor Author

Review round 2: all five fixes pushed, and an answer to each of D1-D5

The head is now 1aca300, a fast-forward from 6d8d822 with eight new commits and no rebase (main is still 985a386). The PR body is updated to match.

Fixes

item commits what changed how it was checked
F1: CI never showed the witness tests 6b0fae3, 8c102f8 A new unit-job step re-runs the 11 engine/iouring witness tests by name. It uses -v, an anchored -run and CELERIS_REQUIRE_IOURING_WORKERS=1, which newTestRing and TestWorkersShareTheHandoffLossWitnesses now honour. The tally requires === RUN Name = 11, --- PASS: Name ( = 11 and no --- SKIP line. The step's block was extracted verbatim and run as GitHub runs shell: bash. The committed tree exits 0 (11/11/11/0). Each copy with one test source edited exits 1: a test deleted, renamed, or calling t.Skip. With io_uring refused by seccomp, the 5 ring-backed tests FAIL ("forbids skipping"). CI tally below.
F2: StaleRecvDataClosed called benign 5d76a8e, b00117f The field doc, the stats doc and the test comment now say Closed can hold a live client's request. After hijackConn the socket lives on under the hijacker's net.Conn. A recv that reaches the kernel after a close resolves the fd number, which a new connection may hold by then. The PR body table says the same. Argued from reading: hijackConn registers through noteClosedInflight, then net.FileConn dups the fd and the original is closed.
F3: recvArmed never tested alone f20fe56 The single-term states now differ from "nothing in flight" in exactly one term: recv armed with kernelInflight 0, a kernel op only, a notification only. A fifth state keeps the usual recv armed and counted. New mutants m18, m19 and m20 each drop one term, and all three are KILLED. With 6d8d822's test file, m18 survives. With the new file it fails: recv armed only: TransplantHandoffInFlight = 0 after this hand-off, want 1.
F4: the 32-byte claim had no record 859eacd, 1aca300 TestClosedOpsEntryStaysThirtyTwoBytes (linux/amd64 and arm64) pins size 32, handoff at offset 4 and conns at offset 8. The comment states the 32-bit growth. m22 (flag after conns) is KILLED. Compile-time unsafe.Sizeof assertions compile only at 32 for the branch on amd64/arm64 and 20 on 386/arm. The base compiles at 32 and 16. Every other asserted size fails to compile.
F5: the buckets are not exhaustive 5d76a8e The field doc and the stats doc now name the generation collision. A stale data CQE whose (fd, gen) matches the current occupant takes staleConnCQE's live branch and none of the three counts it. It needs the process-wide 32-bit connGenSeq to wrap while the op is in flight. It predates this PR. Argued from reading.

I kept all three terms of the in-flight predicate rather than simplifying to kernelInflight != 0, and I agree with the second review as far as it goes. With consistent accounting, every recv arm counts itself, and a SEND_ZC keeps its count until the NOTIF that clears zcNotifPending. But the generation-collision misroute in staleConnCQE can take kernelInflight to 0 while a recv is still armed or a NOTIF is still pending. The three terms are also the exact condition PR 2 refuses a hand-off on.

Two more fixes came out of this round. TestStaleRecvDataCountsEachMultishotCompletion (83a7a4b) pins the F_MORE path from D5, and mutant m23 (count only terminal CQEs) is KILLED. golangci-lint also flagged my first size test (unused fields in a mirror struct), and 1aca300 fixed it.

Mutation now stands at 23/23 KILLED, with the unmutated control at 11/11 and 3/3. golangci-lint 2.13.2 reports 0 issues on linux/amd64 and arm64. actionlint 1.7.12 with shellcheck and zizmor 1.30.0 are clean. go vet and cross-compilation pass on linux/amd64 and arm64, plus 386 and arm for ./engine/iouring. Every run was in Docker, one container at a time.

D1: the behaviour gate was registered weaker than the DECISION

That is right, and I should have said it in the PR. DECISION step 1 asks for "verdict rates within the A/A floor of the unchanged base". I registered Fisher's exact p ≥ 0.05 as the gate and kept the A/A floor only as calibration. Under the DECISION's criterion, campaign V failed TestBidirectionalFlap: |rate M − rate B| = 0.188, above V's largest same-bytes A/A difference, 0.125.

The evidence against a real effect:

TestBidirectionalFlap FAILs base (M) branch (B) M − B
V 14/16 11/16 +0.188
V2 (pre-registered replication on the same bytes, seeded-shuffle order) 11/16 15/16 −0.250
pooled 25/32 26/32 −0.031, Fisher p = 1
  • The lean reversed on identical bytes.
  • The same binary moved by 0.25 between V and V2 (the branch went from 11/16 to 15/16). V's 0.125 floor came from two 8-run halves and was too small to describe the spread. That comparison is post hoc, not registered.
  • TestReverseTransplant stayed within V's own A/A floor (0.312 against 0.500). V2 gave 13 vs 13, and the pooled result was 28 vs 23, p = 0.21.

From reading the code, the PR adds no lock, allocation or reordering on any path:

  • noteHandoffInFlight reads three fields and does one atomic add. It sits before cancelConnOps and moves nothing.
  • noteHandedOffInflight is the old noteClosedInflight call, at the same position, plus one map read and one bool store.
  • noteStaleRecvData does one map read and one atomic add, and runs only for stale CQEs.

The per-request path is untouched.

D2: G2 cannot tell "count armed" from "count every"

Agreed for the integration runs. The distinction is carried by:

  • the quiet-hand-off arms in TestTransplantHandoffInFlightCountsTryTransplant ("nothing in flight", want 0) and TestTransplantHandoffInFlightCountsFinishAsyncTransplant (a quiet async connection, want 0);
  • the new mutant m21, which replaces the predicate with true. It is KILLED by both quiet arms.

An integration arm with quiet hand-offs cannot be built cheaply on this engine. Every hand-off here is made with an op in flight:

  • 0 of 2,205 probe-recorded hand-offs were quiet in the D3 campaign below (1,927 sync, 278 async), and I registered that before the runs;
  • 0 of 2,005 in step 0's T4 probe arm;
  • on main, an idle sync connection is never moved while idle (X1 XM: 64 of 64 stayed for 1.5 s). The connections move when the client resumes, with the recv armed.

The first integration run that can tell the two apart is PR 2's gate. There, TransplantHandoffInFlight must read 0 while TransplantDetached stays above 0.

D3: the T4 shape and the async site

I ran this now, pre-registered: predictions and tools hashed before the build, binary pins added before the first run. Two arms (engine byte for byte, and with the J probe) × two cells × n = 6 gives 24 runs. The shape is Workers=2, 256 keep-alive connections (128 per ring) and 3 promote/revert cycles, at memlock 128 MiB. workers=2 held in 24/24 runs, and all 24 were valid at the first attempt.

sync (TestFlapConnsPerRing) async (same shape, every route async)
StaleRecvDataTransplanted + Unattributed == client read errors; dial = write = 0 12/12; 2,662 == 2,662, loss in 12/12 12/12; 0 == 0, no request lost in any run
TransplantHandoffInFlight == probe hand-offs with an op in flight == kernel-confirmed armed; TransplantDetached == probe hand-offs 6/6 (1,927 at tryTransplant) 6/6 (278 at finishAsyncTransplant)
StaleRecvDataTransplanted == probe-joined stale data CQEs; unjoined 6/6; 0 6/6; 0
  • Every registered prediction held.
  • The join holds across two workers with non-zero loss.
  • The async site's TransplantHandoffInFlight is now join-validated at integration level. Each of its 278 hand-offs was armed, counted and later confirmed by a terminal stale completion. All were cancels, with 0 data.

What stays open is the async site's StaleRecvData* with non-zero loss. No async run lost a request, which I had registered as a possibility (no prediction). I am leaving that to PR 2's failing-first campaign, which runs on main plus this PR. That order is safe for an instruments-only PR, for four reasons:

  1. The classification path does not depend on the site. staleConnCQE → noteStaleRecvData reads only the closedOpsEntry.handoff flag, which both sites set through the same noteHandedOffInflight.
  2. What is specific to the async site is the call itself. It is pinned by TestStaleRecvDataCountsAnAsyncTransplantedConn and mutants m07 and m09, and its integration wiring is the 278 joined hand-offs above.
  3. If async loss were misfiled as Closed, PR 2's gate would still catch it. The DECISION's gate also requires zero client errors, and a tier-2 error with tier-1 at zero is reported as UNEXPLAINED.
  4. The shape that loses async requests on main is known. X1 XM's async idle revert lost 6 requests, and a post-hoc count matched them to 6 stale data CQEs. That cell is in PR 2's plan.

D4: #670 did not pass in all runs

Corrected. TestRampAutoMixedH1H2 (#670) skipped in all 4 of campaign S's 8 MiB adaptive runs, because the up-switch is disabled at a memlock worker ceiling of 1. It passed in all 4 runs at 128 MiB. So S covered #670 at 128 MiB only. The PR-1 report's "passed in all runs" was wrong. The record is corrected in the evidence summary, and the PR body now states the 4/4 skip.

D5: acknowledged

  • F_MORE: never observed. The J probe recorded 1,024 stale recv completions, 491 with data, and none carried F_MORE. It was argued from reading, and is now unit-tested (above).
  • Lint overlap: the lint gate ran on the macOS host, not in a container, from 03:51:55Z to 03:52:51Z. It overlapped the first 9 of the 32 J runs, and the container-count witness could not see a host process. Every J gate held in all 32 runs, those 9 included. This round, every gate ran in a container, one at a time, with no campaign running beside it.
  • Manifest: the round-1 manifest lists ./SHA256SUMS.tmp. This round's manifest is built in the scratchpad, and its script refuses to write a manifest that names itself or a temp file.

CI on 1aca300

All 11 jobs passed on the first attempt (CI run 35424972675, PR run 35424970454), so nothing was re-run. From the real job logs:

  • Unit: the new step ran with shell: /usr/bin/bash --noprofile --norc -e -o pipefail {0}, CELERIS_REQUIRE_IOURING_WORKERS: 1 and memlock (KiB): 8192. Its tally line:
    2026-09-19T05:52:38.1502419Z celeris#657 witness tests: want 11, ran 11, passed 11, SKIP lines 0
    TestWorkersShareTheHandoffLossWitnesses passed at the runner's 8 MiB, so the step needs no raised memlock.
  • Adaptive engine: 75 PASS, 0 FAIL, 0 SKIP (kernel 6.17.0-1022-azure; memlock unlimited; workers=4 ×11, workers=2 ×1). That includes TestReverseTransplant, TestBidirectionalFlap, TestRampAutoMixedH1H2 (adaptive: TestRampAutoMixedH1H2 flakes on the phase-0 migration threshold (949 of the 1024 required) behind a flat 2500 ms sleep #670) and the two new adaptive tests.
  • io_uring init-failure regression: celeris#656 tests passed: 3.

…p comment (celeris#657)

The KNOWN RESIDUAL note in staleConnCQE and cqe.go's layout history still
described a 16-bit, per-connState generation. celeris#470 made it the
32-bit field drawn from the process-wide connGenSeq; both now say so.

The comment at the stale-data counter said every stale recv with data
consumed a request a client is still waiting on. That holds for a
handed-off conn, not for every class; it now points at handoffLossStats.

The witness step's comment said the tests need one ring. One builds two
workers through createWorkers, which skips the Listen path's memlock
ceiling; the comment now says the fit at 8 MiB is measured, and that a
future ENOMEM fails the job on purpose. The step's run block is unchanged.
@FumingPower3925
FumingPower3925 merged commit 5b2e83b into main Sep 19, 2026
12 checks passed
@FumingPower3925
FumingPower3925 deleted the fix/celeris-657-instruments branch September 19, 2026 06:11
FumingPower3925 added a commit that referenced this pull request Sep 20, 2026
…can still resolve its fd (celeris#657) (#681)

PR 2 of 3 for #657. PR 1 (#676) made the loss countable; this stops it.

When io_uring hands a connection to epoll, tryTransplant dup'd the fd,
closed the original and handed the dup over while the connection's next
RECV could still be issued or completed. That recv then resolved the fd
NUMBER, which a later dup may already have given to another connection,
and consumed that connection's request. staleConnCQE dropped the CQE as
stale, so the #624 hand-off ledger balanced while requests went missing.

The rule: never hand a connection off while a read on it can still
resolve its fd.
- R0 gate: refuse the hand-off at both sites while recvArmed,
  kernelInflight != 0 or a SEND_ZC notification is pending.
- REAP: cancel an armed recv first, with a reported cancel under its own
  tag and a per-connection count; the -ECANCELED is routed before the
  generic error branch, and the hand-off re-runs. A miss is retried; any
  other failure is counted and never retried, and never followed by a
  hand-off. A connection whose dup failed is not reaped again until it
  serves another request.
- HOLD: a response served while a drain is set flushes unlinked and does
  not arm the next recv, through one helper for every tail. Every hold is
  released on every path, with a checkTimeouts rescue that must stay 0.
- One owner per hand-off, and pending SQEs are submitted before a worker
  parks.

IORING_ASYNC_CANCEL flags are 5.19+. A cached startup probe detects them;
where they are missing the reap is off (counted), never retried, and no
connection is handed off with a recv armed. Measured on Ubuntu 5.15.0-191,
5.19.0-50 and 6.8.0-138 under QEMU. The pre-existing cancel breakage on
5.10-5.18 is #682.

EngineMetrics gains TransplantHeld, TransplantReaps, TransplantReapMisses,
TransplantReapFailed, TransplantReapUnsupported, TransplantHoldRescued and
TransplantClaimDeferred; TransplantDoubleClaim now counts only a real
second claim, so a release gate of 0 is meaningful.

Measured, one container per observation, pre-registered (304 containers,
prereg 99c7e04c):
- one-worker revert and flap: base FAIL 4/8 and 7/8, fix 0/8 and 0/8.
- T4 (2 workers, 128 conns/ring): base FAIL 8/8 with 1,713 lost requests
  == the W1 counter; fix PASS 8/8.
- switch-stall harness: base 18/60 stall rounds, fix 0/60 (p=1.7e-6).
- W1, W2, DoubleClaim, HoldRescued and ReapFailed are 0 in all 116 fix
  records, while the fix still handed off every connection.
- Every negative control loses: no HOLD/REAP, blind witnesses, no A6, no
  A5. Mutants: 85 killed, 0 survivors.
- Refused dials at a revert, pre-registered on this head: no regression
  (base 7/104, fix 5/104, inside the A/A floor). The listener gap those
  dials fall into is pre-existing and is #683.
- No PASS->FAIL in any package x memlock cell.

CI: the witness step tallies 29 tests with skipping forbidden, a new
one-worker leg requires err=0, and T4 runs with a guarded tally. Each was
proven to fail on a deleted, renamed, skipped or loss-carrying tree, and a
deliberate red run on the runner (35465814564) showed T4 and the err=0 leg
turn red for the planted reason; its revert is green.

Not in this PR: the placement half of #657 (sweep, A2w, A1) is PR 3.
Tracked elsewhere: #682, #683, #684, #685.

Refs #657
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working engine/iouring io_uring engine specifics

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant