You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Filed from the release-readiness sweep of 2026-09-13 under the standing rule that a closure counts only when it is MEASURED. Design produced by a read-only design pass and corrected by two adversarial reviews; the corrections are binding and listed last.
Closure under test
celeris#465 — SEND_ZC default kept ON by PR #489 "pending fabric A/B (~2026-09-08)"; the A/B was never run or recorded.
What the closure claims
Issue CLOSED (milestone v1.6.0) by PR #489 (merged 2026-09-07T17:51Z, merge e468d3b). The PR fixed notifUsageZCCopied (uint32 1<<31, probe.go:74/234), split the loopback probe into a FUNCTIONAL check (copy-fallback on loopback = functional, engine.go:80-95) and kept the production default enabled: resolveSendZCPolicy(functional, "auto"/"") returns functional (probe.go:249-263), i.e. send_zc=true on every host whose kernel accepts the opcode. The PR body, doc.go:9-13 and probe.go:246-247 all say the FINAL default is "pending a measured cluster fabric A/B (SEND vs SEND_ZC on get-json-64k, ws-large-echo, streaming cell, alternating builds, >=3 trials) once the cluster is free (~2026-09-08)". As of 2026-09-13 no benchmark run has toggled CELERIS_IOURING_SEND_ZC (last successful Benchmark Tier = 33256654220, 2026-08-29, pre-PR; every later run cancelled/none), so the perf-affecting default was decided by argument, and the bench that will publish v1.6.0 would run it unmeasured.
What was argued rather than measured
(1) A loopback self-probe can never observe the NIC (kernel always copies on local delivery), NETIF_F_SG gating is unsound, so keep a functional probe + an env knob (CELERIS_IOURING_SEND_ZC=on|off|auto) and decide the default later by A/B — sound. (2) "Blast radius is narrow: unlinked sends >= sendZCMinBytes, essentially get-json-64k and ws-large-echo" — PARTLY WRONG on both facts: get-json-64k is not in the grid (probatorium scenarios/static.go:144-153 cut all 8k/16k/64k rows as wire-bound; the /json-64k handler exists in servers/celeris/server.go:172 but is unscheduled), and the PR's "sendZCMinBytes = 32768" is stale — consts.go:114 is 4096 (commit ce1073c, already at tag v1.5.8). (3) Never stated, but it determines the whole A/B: on the default bench configuration SEND_ZC is STRUCTURALLY UNREACHABLE for HTTP. Multishot recv is opt-in (doc.go:15-16) so w.bufRing == nil, and every H1 response then goes through flushSendLink (worker.go:1906, 2162, 2338) which always emits a plain linked SEND (worker.go:4034-4046, comment "never ZC"); bodies staged by writeBody go WRITEV (worker.go:3922-3950); h2 frames are written inline with unix.Write (worker.go:3441) and only the EAGAIN/partial remainder reaches the ring. The only unlinked ring SEND in the grid is the detached WS/SSE path: inline unix.Write on the dispatch goroutine (worker.go:1567-1575), remainder -> detachQueue -> drainDetachQueue -> flushSend (unlinked) -> prepSendSQE -> useSendZC(sendZC, false, n>=4096) (worker.go:3896-3991). ws-echo/ws-hub/sse-fanout frames are 256 B (< 4096), so ws-large-echo (64 KiB, scenarios/streaming.go:132) is the ONE cell where the default matters. (4) There is already an un-attributed signal exactly there: the five byte-identical published v1.5.8 amd64 runs (docs/results/v1.5.8/{20260716,0722,0729,0805,0829}/x86_64/summary.json; ZC was ON via the buggy probe, threshold 4096) show every io_uring column on ws-large-echo at 95.6-96.8 % SUT CPU (resources.summary.mean_cpu_pct) for 26.6-29.7 K msg/s (1.76-1.95 GB/s one-way), while epoll delivers the same 26.0-29.5 K at 53-58 % and std 34.8-35.4 K at 56-62 %. io_uring spends ~1.7x the CPU per byte of epoll on precisely the ZC-eligible cell; whether SEND_ZC (extra NOTIF CQE per send, buffer held across DMA, possible kernel copy-fallback via skb_copy_ubufs) is that cost is the unanswered question. The 2026-09-04 review already scored the default choice as "unmeasured perf-affecting" (memory v160-bug-plan-review); the 2026-09-13 release sweep flagged #465 as argued-not-measured.
Measurement
Observable
Per cell, from the harness artefacts already produced for every bench cell: (a) saturation_mode_rps of ws-large-echo (msg/s; 1 msg = one 64 KiB frame out + echo back, loadgen ws.go:33/138-153), (b) resources.summary.mean_cpu_pct = SUT host-wide mpstat 'all'-row busy % at 1 Hz (report/resources.go:100-143, sampler = run_bench_cell.yml:384-389), (c) derived DECISION METRIC cpu_per_GB = mean_cpu_pct / (rps * 65536 B) i.e. SUT CPU per delivered byte, (d) rated_mode_p99_at_target_rps / latency_at_slo for ws-large-echo as secondary. Arm identity is proven per cell by the engine log line 'io_uring engine selected ... send_zc=true|false' (engine.go:137-145, slog.Default -> stderr -> {{cell_dir}}/server.log, run_bench_cell.yml:271). Controls in the SAME dispatch: ws-echo (256 B < sendZCMinBytes, identical detached write path, useSendZC always false) and get-json (linked SEND only). Optional kernel-side observables (root on the SUT, outside the harness): count of io_uring submissions with opcode 47 (IORING_OP_SEND_ZC, consts.go:45) vs 26 (SEND) during the cell, and kprobe:skb_copy_ubufs hit count (= number of ZC sends the kernel COPIED because the device/bond cannot take user frags).
Where it runs
probatorium Benchmark Tier on the cluster (benchmark-tier.yml), profile=fast (35 s/10 s, saturation-only, BENCH_PUBLISH forced 0 at benchmark-tier.yml:154 so nothing reaches docs), competitors=celeris-iouring-h1-async, deploy_competitors=celeris, ONE ARCH PER DISPATCH: target=msa2-server (amd64, 2x10G LACP bond, kernel 7.0.0-30, memlock unlimited so the ENOMEM fallback at worker.go:2434/2521 cannot silently disable ZC mid-run) first, then target=msr1 (arm64, RTL8127/r8127 DKMS NIC) as a second series. NOT target=both: the 20260829 parallel-arch run depressed ws-large-echo ~8 % on EVERY amd64 column (27.6-29.7 K -> 26.0-26.9 K) because this cell runs near the loadgen bond's line rate and the two arch passes share msa2-client. Bench SUT is pinned to celeris main ededb6c (servers/celeris/go.mod:6), i.e. the v1.6.0 candidate.
Procedure
Land the two small harness passthroughs in code_changes_needed (env + cells), on main, before dispatching. 2. Wait for the cluster group to drain (24 h soak 34736980002 ends ~2026-09-14T04:15Z; nightly 34744788978 is pending behind it; matrix-tier-cluster serializes). 3. Dispatch SIX runs on amd64 in the order ON, OFF, OFF, ON, ON, OFF (A B B A A B cancels monotone drift), each: gh workflow run benchmark-tier.yml --repo goceleris/probatorium -f profile=fast -f target=msa2-server -f competitors=celeris-iouring-h1-async -f deploy_competitors=celeris -f publish=false -f cells='ws-large-echo/,ws-echo/,get-json/' -f sut_env='CELERIS_IOURING_SEND_ZC=on' (or =off). Use the explicit 'on', not 'auto', so the ON arm is not hostage to the functional probe. 4. For each run read benchmark-tier-results-<run_id> (results.json + 00-celeris-iouring-h1-async/server.log + run0//.json + cpu.log): assert server.log contains send_zc=true for ON and send_zc=false for OFF (a mismatch voids the run); extract rps, mean_cpu_pct, p99 for the 3 cells. 5. Form the three adjacent pairs (1-2, 3-4, 5-6), compute Delta = OFF - ON for cpu_per_GB and rps on ws-large-echo, and the same deltas on ws-echo and get-json. 6. Decision rule: WIN for ZC if cpu_per_GB(ON) < cpu_per_GB(OFF) by >= 5 % in 3/3 pairs AND rps(ON) >= rps(OFF) - 3 %; LOSS if cpu_per_GB(ON) > cpu_per_GB(OFF) by >= 5 % in 3/3 pairs OR rps(ON) < rps(OFF) by >= 3 % in 3/3 pairs; NOISE if neither (or the controls move by as much as the treatment cell, which flags run-level drift). If the three pairs disagree in sign, add one more pair (2 dispatches) before calling it noise. 7. Optional but decisive for interpretation, during ONE ON-arm ws-large-echo cell on msa2-server as root: bpftrace -e 'tracepoint:io_uring:io_uring_submit_req /args->opcode==47/ {@zc=count()} tracepoint:io_uring:io_uring_submit_req /args->opcode==26/ {@send=count()} kprobe:skb_copy_ubufs {@copied=count()}' for 30 s; plus ethtool -k bond0 and each slave | grep scatter-gather. 8. Repeat steps 3-6 with target=msr1 for arm64 (its ws-large-echo is CPU-bound at 11 K msg/s / 95.6 % CPU, so rps is the primary metric there). 9. If LOSS: change resolveSendZCPolicy 'auto' to return false (probe.go:258-259), re-run one confirming pair, repin probatorium, THEN the release bench, so the flip lands at the version boundary the issue itself asked for. If WIN or NOISE: record the numbers in io_uring: probeSendZC can never detect copy-fallback — SEND_ZC enabled on every host regardless of NIC #465 and leave the default ON.
Expected if the claim is TRUE
Claim under test = 'SEND_ZC on the real fabric is a win, so keeping the default ON is right'. Expected: in every one of the 3 amd64 pairs the ON arm shows lower cpu_per_GB on ws-large-echo by >= 5 % (e.g. mean_cpu_pct drops from the ~96 % plateau of the v1.5.8 data toward epoll's 53-58 % at the same 27-29.7 K msg/s, or rps rises above ~29.7 K at the same CPU); bpftrace shows opcode-47 submissions in the thousands per second during the cell with skb_copy_ubufs ~0 (true zero-copy on the bond); the controls ws-echo and get-json differ between arms by <= 1.6 % rps and <= 0.6 CPU points. On arm64 ON rps >= OFF rps + 3 %.
Expected if the claim is FALSE
Two distinct falsifications. LOSS: the OFF arm has lower cpu_per_GB by >= 5 % and/or higher rps by >= 3 % in 3/3 pairs — most plausibly OFF brings the io_uring column's ~96 % CPU down toward epoll's ~55-58 % at equal throughput (the NOTIF-per-send + buffer-hold cost, or a copy-fallback where bpftrace shows skb_copy_ubufs count ~= opcode-47 count, meaning the kernel copies anyway and ZC only adds a second CQE). NOISE: |Delta cpu_per_GB| < 5 % and |Delta rps| < 3 % in every pair with inconsistent sign, controls moving as much as the treatment; then the default is immaterial for the published grid and #465's residual question is closed as 'no measurable effect on this fabric', which is also a real result. Either way the ~1.7x CPU-per-byte gap between io_uring and epoll on this cell becomes attributable (to ZC, or explicitly NOT to ZC -> a separate issue on the detached-write funnel).
Negative controls
Three built into every dispatch: (1) ws-echo — same detached inline-write/detachQueue path, 256 B frames < 4096 so useSendZC is false (send_zc_gate_test.go:24-25 pins this) — must show no arm difference; (2) get-json — linked SEND only, ZC unreachable — must show no arm difference; if either control moves as much as ws-large-echo the delta is run-level drift, not ZC. (3) Arm verification: server.log must read send_zc=false in the OFF arm; the bpftrace opcode-47 count must be 0 in the OFF arm and > 0 in the ON arm (proves the env reached a root-launched SUT and that the cell actually exercises the branch — without this a 'noise' verdict could just mean the toggle never applied or the branch never fired). Existing-data sanity anchor: the ON arm at profile=fast should reproduce the v1.5.8 fast-run band (28.8-29.7 K msg/s, 96.3-96.8 % CPU) within the v1.6.0 engine changes; a large departure from that band in BOTH arms indicates something else moved (#572/#574/#525 etc.), to be reported, not folded into the ZC verdict.
Cost
Per dispatch with the cells passthrough (3 cells x 1 column x 1 arch): setup runner ~5-10 min + mage Deploy (Go build of servers/celeris only) ~5-10 min + 3 cells x ~2 min (budget model 10+35+5+12 = 62 s/cell, ~2x observed under-count) + merge/upload/cleanup/teardown ~10-15 min = ~35-45 min, i.e. ~4-4.5 h cluster wall-clock for the six amd64 dispatches; the same again for arm64 (msr1) = ~9 h total for both arches, versus ~74 h for the release bench. WITHOUT the cells passthrough (column-scoped only via the existing competitors input) each dispatch runs the column's ~29 capability-gated scenarios (~1 h at fast + overhead) = ~1.5 h/dispatch, ~9 h per arch, ~18 h both — the controls then come for free. Engineering: ~1-2 h for the passthroughs + a unit test. Statistical cost: three pairs give a 3/3 sign requirement (one-sided p = 0.125 on sign alone) — the decision rests on the magnitude thresholds being 3x the measured same-window spread (rps <= 1.6 %, CPU <= 0.6 pt over five byte-identical v1.5.8 runs), not on the sign test.
Not measurable, and why
(1) Whether the NIC/bond actually zero-copies cannot be observed from celeris: production SQEs never set IORING_SEND_ZC_REPORT_USAGE (sqe.go:175-186 sets only opcode/fd/buf), so NOTIF CQEs carry no ZC_COPIED bit; only the out-of-harness kernel probe (skb_copy_ubufs / tracepoint opcode filter, root on the SUT) can tell copy-fallback from true ZC, and the harness cannot record it. (2) The fraction of ws-large-echo bytes that go through SEND_ZC vs the inline unix.Write fast path (worker.go:1567) is not counted anywhere — EngineMetrics (engine.go:335-351) has bytesWritten but no ZC counter — so a NOISE verdict could mean 'ZC rarely fires' rather than 'ZC is free'; the bpftrace step or a small celeris counter is the only way to separate these. (3) HTTP/1.1 and h2 large responses are structurally ZC-unreachable in the bench (linked SEND, WRITEV bodies, inline h2 writes), so this A/B says NOTHING about ZC for HTTP; the only configuration where every >= 4 KiB send would be ZC is multishot recv (CELERIS_IOURING_MULTISHOT_RECV=1, bufRing != nil -> flushSend unlinked), which the bench does not run — a second A/B dimension deliberately out of scope. (4) Fixed-file mode (off by default, celeris#541/#553) and the adaptive columns' post-promotion behaviour are not covered. (5) The CI memlock-limited ENOMEM fallback path (worker.go:2434, 2521; 8 MB memlock on GitHub runners) is not exercised on the cluster (memlock unlimited). (6) Cross-dispatch rps is confounded by window length (35 s vs 90 s) and by any parallel arch pass, so only within-pair comparisons at identical profile/target are valid; a single ON run compared against the historical v1.5.8 numbers is NOT a measurement. (7) mean_cpu_pct averages the whole cell-guard window including idle edges (dilutes both arms equally) and is host-wide, so it cannot attribute CPU to the SUT process vs kernel softirq — a perf split is out of harness.
Code changes needed
Required (probatorium, harness only; no celeris change needed to run the A/B): (1) SUT env passthrough — benchmark-tier.yml: add workflow_dispatch input sut_env (KEY=VALUE, default "") -> job env BENCH_SUT_ENV; mage_bench.go Bench() buildArgs (mage_bench.go:451-483): when BENCH_SUT_ENV is set, append --extra-vars with a JSON dict {"bench_sut_env": {KEY: VALUE}}; ansible/tasks/run_bench_cell.yml 'start competitor server in background' (lines 199-272): turn the hard-coded environment: mapping (207-235) into environment: "{{ sut_base_env | combine(bench_sut_env | default({})) }}" and also log the merged env into server.log (or echo it in the launch shell) so the arm is recorded next to the send_zc= line. (2) Cell-scope passthrough — benchmark-tier.yml: input cells (default "") -> env BENCH_CELLS; mage_tier.go setBenchEnvFromProfile (lines 218-244): honour a pre-set BENCH_CELLS the same way it already honours BENCH_DURATION/BENCH_WARMUP instead of unconditionally os.Setenv("BENCH_CELLS", budget.CellsGlob(p)); FitWithin still projects p.Cells=813 so it passes trivially — acceptable for a scoped run, note it in the log. Unit tests: a table test for the merged extra-vars and one for the BENCH_CELLS precedence. Optional but recommended for interpretability: (3) celeris — two monotonic counters on Worker (zcSendsSubmitted in prepSendSQE's ZC branch worker.go:3973-3977, zcNotifs in handleSend worker.go:2370) surfaced through EngineMetrics (engine.go:335-351) and thus /debug/vars, which the bench observer already scrapes for celeris cells (run_bench_cell.yml:421) — this makes the 'did ZC actually fire' question harness-recordable instead of bpftrace-only; small, additive, covered by extending send_zc_gate_test.go. (4) Optional: record bench_sut_env in the merged Document's environment (mage_bench.go:1788-1806) so an arm is identifiable from results.json alone. If the verdict is LOSS: celeris probe.go:258-259 'auto' -> return false (default OFF; explicit on still honoured), doc.go:9-13 wording, and a repin — that is the follow-up, not part of the measurement.
Corrections from the adversarial review (binding)
Review 1: refuted
Why: Two independent defects each make the proposed A/B unable to separate "SEND_ZC is a win/loss" from "SEND_ZC never mattered", and the design's decision rule then converts the guaranteed NOISE outcome into a verdict.
The decision metric does not exist per cell. resources.summary.mean_cpu_pct is the mean of ONE mpstat run over the WHOLE competitor column: in the cluster harness a "cell" is (run_index, competitor_slug) (probatorium ansible/tasks/run_bench_cell.yml:2-6), the server is launched once and mpstat -P ALL 1 {{cell_guard_seconds}} runs once per column (lines 384-389), and aggregatePerCellResults stamps that single aggregate onto every scenario the runner expanded ("The same aggregate applies to every scenario ... since the observer scopes to the whole cell process", mage_bench.go:996-1004). Verified in the published data the design cites: in docs/results/v1.5.8/20260805/x86_64/summary.json every celeris column has 27 scenarios, exactly 1 distinct resources summary, byte-identical series across ws-large-echo and get-json, spanning 1334 s (8024 s in the 20260829 full run). So the motivating "existing signal" — io_uring 95.6-96.8 % vs epoll 53-58 % "on precisely the ZC-eligible cell" — is the column-wide average over churn-close, driver-*, get-json, etc., and says nothing about ws-large-echo. Under the proposed 3-cell scoping the treatment cell and BOTH controls would share one CPU scalar: the control test ("ws-echo/get-json must move <= 0.6 CPU pt") is unevaluable, and any treatment delta is diluted by two control cells plus cooldown/idle gaps. cpu_per_GB as specified cannot discriminate ON from OFF. rps cannot rescue it on amd64: at 29.7 K msg/s x 64 KiB x 2 directions the cell is near the 2x10G bond's line rate (std reaches 35 K), which is why the design itself leans on CPU.
Under the bench configuration the ZC branch almost never carries bytes, so the toggle is nearly inert. Detached WS egress first tries an inline unix.Write(cs.fd, cs.writeBuf) on the dispatch goroutine and only hands the PARTIAL remainder to the ring (celeris engine/iouring/worker.go:1567-1583 guarded closure; ring path = drainDetachQueue -> flushSend -> prepSendSQE -> useSendZC, worker.go:3896-3991). loadgen's ws-large-echo is strictly ping-pong — exchange writes one frame then blocks until its echo arrives (loadgen ws.go:138-163) — so each conn has at most one 64 KiB frame in flight and an empty send queue when the echo is written; the SUT runs kernel-default autotuned tcp_wmem (bench_tuning.yml touches no wmem/rmem sysctl). In steady state the whole frame goes inline; SEND_ZC only fires for the rare short-write remainder that is also >= 4096 B. The inline path landed in 0af8831 (2026-07-03) and is inside v1.5.7 and v1.5.8, so even the cited 96 % number was produced with ZC essentially unreachable. The design concedes this in not_measurable(2) but its decision rule maps NOISE to "leave the default ON, close io_uring: probeSendZC can never detect copy-fallback — SEND_ZC enabled on every host regardless of NIC #465 as no measurable effect" — the result the run would produce whether a ZC send costs 0 % or 50 % more CPU than a SEND, because there is no ZC-fired count to distinguish "ZC is free" from "ZC never executed". The bpftrace count is optional and one-arm-only, and the celeris ZC counters are optional.
Minor: neither the sut_env nor the cells workflow input exists yet, so step 3's command line cannot run today (the design acknowledges this). The arm-identity check is sound: profile.SendZC = enabled precedes SelectTier(profile) (engine.go:97-103), so the logged send_zc= is the effective post-env value.
Change to the design: Make both legs discriminating before spending cluster time:
A. Per-cell CPU. Never use the summary scalar. Compute per-scenario CPU by windowing the raw 1 Hz cpu.log (wall-clock rows; report/resources.go ParseMPStat already returns the series) against the runner's per-scenario started_at/completed_at recorded in results.json (cmd/runner/main.go:1093-1096), trimming the warmup window; or have the harness restart mpstat per scenario. Add SUT-process CPU (utime+stime from /proc//stat, the observer already owns the PID: run_bench_cell.yml:420) and mpstat's %soft so a ZC cost (kernel NOTIF/skb_copy_ubufs work) is attributable rather than host-wide. Define cpu_per_GB and the control comparisons on these windowed numbers only. Before dispatching anything, apply the same windowing to a retained v1.5.8/20260829 run artefact to learn whether io_uring's per-cell CPU on ws-large-echo actually differs from epoll's — the current 96 % vs 55 % motivation is a column average and must be retracted.
B. Make ZC carry the bytes, and prove it. (1) Promote the celeris counters from optional to REQUIRED: zcSendsSubmitted, zcNotifs, inlineBytes (the unix.Write fast path) and ringBytes (flushSend/WRITEV) on Worker, surfaced via EngineMetrics -> /debug/vars, scraped by the observer per cell; a verdict is only valid if ringBytes/(inlineBytes+ringBytes) >= ~50 % and zcSendsSubmitted is in the thousands/s in the ON arm and 0 in the OFF arm; otherwise the run has answered a different question. (2) Because the production ping-pong path takes the inline fast path, run a 2x2: {SEND_ZC on/off} x {inline egress on/off}. Add a celeris env knob (e.g. CELERIS_IOURING_INLINE_EGRESS=off) that skips the unix.Write fast path so every detached >= 4 KiB send reaches prepSendSQE; the inline-off pair is the only one that measures SEND vs SEND_ZC on the fabric, and the inline-on pair measures whether the default matters in production (expected: it does not fire — which is then a real, counted result, not an inferred one). Alternatively (no celeris knob) add a large-frame fan-out cell (ws-hub-broadcast with 64 KiB payload, the shape celeris's own broadcast_egress_correctness_linux_test.go uses) where per-conn queues fill and the ring path genuinely carries the data. (3) Keep bpftrace opcode-47/26 + kprobe:skb_copy_ubufs, but run it in BOTH arms of at least one pair and make it a precondition (opcode-47 > 0 ON, == 0 OFF), plus ethtool -k bond0 and slaves | grep scatter-gather, so copy-fallback vs true ZC is on record.
C. Keep ABBAAB ordering, one arch per dispatch, server.log send_zc= arm check, and the harness passthroughs (sut_env, cells). Restate the decision rule on the windowed per-cell metrics and the ZC-fired precondition: WIN/LOSS thresholds as before; NOISE is only a verdict when the ON arm demonstrably routed >= 50 % of ws-large-echo bytes through SEND_ZC; a NOISE result with a low ZC fraction closes #465's production question ("the default is immaterial because the branch does not fire on the echo path") but leaves the fabric question open for the inline-off pair.
Review 2: refuted
Why: The design's DECISION METRIC (cpu_per_GB = mean_cpu_pct / bytes on ws-large-echo) and the premise that motivates the whole A/B ("io_uring at 95.6-96.8 % SUT CPU on precisely the ZC-eligible cell vs epoll 53-58 %") both rest on resources.summary.mean_cpu_pct being a per-(column, scenario) number. It is not. In the harness a "cell" is one COLUMN's whole runner pass: bench.yml:266-282 builds the schedule as [(0, competitor)...], run_bench_cell.yml:42 sets cell_dir = <run_dir>/00-<competitor>, the mpstat sampler (run_bench_cell.yml:384-389) and the observer (:404-426) are started ONCE per that dir for cell_guard_seconds, and the single runner invocation (:514-525) runs every scenario matched by -cells inside it. The merge then copies the one aggregate onto every scenario — mage_bench.go:996-1004, in its own words: "The same aggregate applies to every scenario the runner expanded in this cell, since the observer scopes to the whole cell process." The published data proves it: in all six v1.5.8 summary.json files every column has exactly ONE distinct mean_cpu_pct across its 27-29 scenarios (e.g. celeris-iouring-h1-async 20260829 amd64: churn-close, get-json, ws-echo, ws-large-echo ... all 96.283 %; epoll-h1-async all 53.31 %), and the "ws-large-echo" series spans 1334 s (fast) / 8024-8850 s (90 s runs) — the whole column — not one ~35-90 s cell. So the "1.7x CPU per byte on the ZC-eligible cell" is an artefact of reading a column-wide mean (dominated by 27 other scenarios incl. get-json at 1 M rps and churn-close) as a per-cell value; the io_uring-vs-epoll CPU gap is column-wide (consistent with worker.go:1252 adaptiveTimeout()==0 whenever dirtyHead != nil, i.e. a busy loop under any pending send) and says nothing specific about SEND_ZC. Consequences for the proposed measurement as written: (a) with the planned cells='ws-large-echo/*,ws-echo/*,get-json/*' passthrough, all three scenarios run in ONE column pass and share ONE cpu.log, so the treatment cell and both "negative controls" would report the IDENTICAL mean_cpu_pct by construction — the CPU control is vacuous and cannot flag drift; (b) that shared number would be the average over get-json (CPU-saturated at ~1 M rps regardless of ZC), ws-echo, ws-large-echo and the idle/warmup edges, so a real ZC-driven CPU change on ws-large-echo (which contributes ~1/3 of the window) is diluted at least 3x before the 5 % threshold is applied, biasing the rule toward NOISE — and the design explicitly closes #465 on NOISE as "no measurable effect on this fabric"; (c) the 5 %/3 % thresholds were calibrated against "same-window spread ... CPU <= 0.6 pt over five v1.5.8 runs", which is the spread of a column-wide mean over ~27 scenarios, not the per-cell noise floor, so the threshold itself is mis-sized. Secondary, non-load-bearing errors: the "expected_if_true" anchor "mean_cpu_pct drops from the ~96 % plateau toward epoll's 53-58 %" compares column-wide means; and resources.series cpu_pct points are the same column-wide 60-point downsample. What is NOT refuted: saturation_mode_rps on ws-large-echo is a genuine per-scenario observable (loadgen ws.go:138-153 is a synchronous write-then-read per conn, so msg/s tracks SUT turnaround; std reaching 34.8-35.4 K on the same fabric shows io_uring's 26-29.7 K is not wire-capped), the send_zc=true|false log-line arm check, the code-path analysis (inline unix.Write fast path at worker.go:1567 with ring SEND -> useSendZC(n>=4096) only on the EAGAIN/partial remainder; flushSendLink 'never ZC' for H1), and the ABBAAB ordering. The design would produce numbers, but its CPU-per-byte verdict would be about the column, not the cell.
Change to the design: Make the CPU observable per-scenario before running anything, and re-base the decision rule on it. Concretely: (1) In the merge (mage_bench.go readCellResources / report.SummarizeResources) window the column's cpu.log and observer.sqlite per scenario using the per-scenario started_at/finished_at the runner already writes (cmd/runner/main.go:304/339, report/schema.go:296-297; mpstat rows carry wall-clock times, observer rows carry ts_unix), and attach the sliced ResourceStats to each scenario record instead of the shared aggregate; alternatively, and with zero merge changes, dispatch ONE scenario per column pass (cells='ws-large-echo/*' only, and separate dispatches for the ws-echo and get-json controls) so the column-wide mean IS the cell mean — but then compute the mean over the measurement window only (exclude the warmup/idle edges via the timeseries t in run0//*.json), because the guard window is ~5/4 x budget + 300 s and mostly idle for a 35 s cell. (2) Replace host-wide mpstat as the primary with per-process SUT CPU (utime+stime of the server.pid from /proc, which the observer already sidecars, or pidstat -p) so worker busy-spin/softirq is at least separable; keep mpstat 'all' as secondary. (3) Re-derive the noise floor and the 5 %/3 % thresholds from per-cell repeats (>=3 identical-arm ws-large-echo runs), not from the column-wide spread. (4) Because the io_uring worker spins whenever dirtyHead != nil (adaptiveTimeout()==0), report and reason on rps as the primary decision metric on both arches (it is per-scenario and not wire-capped), with per-scenario CPU as the secondary; do not close #465 on a CPU NOISE verdict unless the per-scenario CPU is shown to move with a known perturbation (e.g. the ws-echo vs ws-large-echo difference) so the metric is proven sensitive. (5) Make the ZC-fired counter (zcSendsSubmitted / zcNotifs via EngineMetrics -> /debug/vars, scraped by the observer) REQUIRED rather than optional, so NOISE can be distinguished from "ZC never fired" on this path; the bpftrace step can stay optional. (6) Correct the premise in the issue write-up: the 96 % vs 55 % io_uring-vs-epoll CPU gap is column-wide across all ~27 scenarios and belongs to a separate io_uring loop-CPU issue, not to #465.
Filed from the release-readiness sweep of 2026-09-13 under the standing rule that a closure counts only when it is MEASURED. Design produced by a read-only design pass and corrected by two adversarial reviews; the corrections are binding and listed last.
Closure under test
celeris#465 — SEND_ZC default kept ON by PR #489 "pending fabric A/B (~2026-09-08)"; the A/B was never run or recorded.
What the closure claims
Issue CLOSED (milestone v1.6.0) by PR #489 (merged 2026-09-07T17:51Z, merge e468d3b). The PR fixed notifUsageZCCopied (uint32 1<<31, probe.go:74/234), split the loopback probe into a FUNCTIONAL check (copy-fallback on loopback = functional, engine.go:80-95) and kept the production default enabled: resolveSendZCPolicy(functional, "auto"/"") returns functional (probe.go:249-263), i.e. send_zc=true on every host whose kernel accepts the opcode. The PR body, doc.go:9-13 and probe.go:246-247 all say the FINAL default is "pending a measured cluster fabric A/B (SEND vs SEND_ZC on get-json-64k, ws-large-echo, streaming cell, alternating builds, >=3 trials) once the cluster is free (~2026-09-08)". As of 2026-09-13 no benchmark run has toggled CELERIS_IOURING_SEND_ZC (last successful Benchmark Tier = 33256654220, 2026-08-29, pre-PR; every later run cancelled/none), so the perf-affecting default was decided by argument, and the bench that will publish v1.6.0 would run it unmeasured.
What was argued rather than measured
(1) A loopback self-probe can never observe the NIC (kernel always copies on local delivery), NETIF_F_SG gating is unsound, so keep a functional probe + an env knob (CELERIS_IOURING_SEND_ZC=on|off|auto) and decide the default later by A/B — sound. (2) "Blast radius is narrow: unlinked sends >= sendZCMinBytes, essentially get-json-64k and ws-large-echo" — PARTLY WRONG on both facts: get-json-64k is not in the grid (probatorium scenarios/static.go:144-153 cut all 8k/16k/64k rows as wire-bound; the /json-64k handler exists in servers/celeris/server.go:172 but is unscheduled), and the PR's "sendZCMinBytes = 32768" is stale — consts.go:114 is 4096 (commit ce1073c, already at tag v1.5.8). (3) Never stated, but it determines the whole A/B: on the default bench configuration SEND_ZC is STRUCTURALLY UNREACHABLE for HTTP. Multishot recv is opt-in (doc.go:15-16) so w.bufRing == nil, and every H1 response then goes through flushSendLink (worker.go:1906, 2162, 2338) which always emits a plain linked SEND (worker.go:4034-4046, comment "never ZC"); bodies staged by writeBody go WRITEV (worker.go:3922-3950); h2 frames are written inline with unix.Write (worker.go:3441) and only the EAGAIN/partial remainder reaches the ring. The only unlinked ring SEND in the grid is the detached WS/SSE path: inline unix.Write on the dispatch goroutine (worker.go:1567-1575), remainder -> detachQueue -> drainDetachQueue -> flushSend (unlinked) -> prepSendSQE -> useSendZC(sendZC, false, n>=4096) (worker.go:3896-3991). ws-echo/ws-hub/sse-fanout frames are 256 B (< 4096), so ws-large-echo (64 KiB, scenarios/streaming.go:132) is the ONE cell where the default matters. (4) There is already an un-attributed signal exactly there: the five byte-identical published v1.5.8 amd64 runs (docs/results/v1.5.8/{20260716,0722,0729,0805,0829}/x86_64/summary.json; ZC was ON via the buggy probe, threshold 4096) show every io_uring column on ws-large-echo at 95.6-96.8 % SUT CPU (resources.summary.mean_cpu_pct) for 26.6-29.7 K msg/s (1.76-1.95 GB/s one-way), while epoll delivers the same 26.0-29.5 K at 53-58 % and std 34.8-35.4 K at 56-62 %. io_uring spends ~1.7x the CPU per byte of epoll on precisely the ZC-eligible cell; whether SEND_ZC (extra NOTIF CQE per send, buffer held across DMA, possible kernel copy-fallback via skb_copy_ubufs) is that cost is the unanswered question. The 2026-09-04 review already scored the default choice as "unmeasured perf-affecting" (memory v160-bug-plan-review); the 2026-09-13 release sweep flagged #465 as argued-not-measured.
Measurement
Observable
Per cell, from the harness artefacts already produced for every bench cell: (a) saturation_mode_rps of ws-large-echo (msg/s; 1 msg = one 64 KiB frame out + echo back, loadgen ws.go:33/138-153), (b) resources.summary.mean_cpu_pct = SUT host-wide mpstat 'all'-row busy % at 1 Hz (report/resources.go:100-143, sampler = run_bench_cell.yml:384-389), (c) derived DECISION METRIC cpu_per_GB = mean_cpu_pct / (rps * 65536 B) i.e. SUT CPU per delivered byte, (d) rated_mode_p99_at_target_rps / latency_at_slo for ws-large-echo as secondary. Arm identity is proven per cell by the engine log line 'io_uring engine selected ... send_zc=true|false' (engine.go:137-145, slog.Default -> stderr -> {{cell_dir}}/server.log, run_bench_cell.yml:271). Controls in the SAME dispatch: ws-echo (256 B < sendZCMinBytes, identical detached write path, useSendZC always false) and get-json (linked SEND only). Optional kernel-side observables (root on the SUT, outside the harness): count of io_uring submissions with opcode 47 (IORING_OP_SEND_ZC, consts.go:45) vs 26 (SEND) during the cell, and kprobe:skb_copy_ubufs hit count (= number of ZC sends the kernel COPIED because the device/bond cannot take user frags).
Where it runs
probatorium Benchmark Tier on the cluster (benchmark-tier.yml), profile=fast (35 s/10 s, saturation-only, BENCH_PUBLISH forced 0 at benchmark-tier.yml:154 so nothing reaches docs), competitors=celeris-iouring-h1-async, deploy_competitors=celeris, ONE ARCH PER DISPATCH: target=msa2-server (amd64, 2x10G LACP bond, kernel 7.0.0-30, memlock unlimited so the ENOMEM fallback at worker.go:2434/2521 cannot silently disable ZC mid-run) first, then target=msr1 (arm64, RTL8127/r8127 DKMS NIC) as a second series. NOT target=both: the 20260829 parallel-arch run depressed ws-large-echo ~8 % on EVERY amd64 column (27.6-29.7 K -> 26.0-26.9 K) because this cell runs near the loadgen bond's line rate and the two arch passes share msa2-client. Bench SUT is pinned to celeris main ededb6c (servers/celeris/go.mod:6), i.e. the v1.6.0 candidate.
Procedure
Expected if the claim is TRUE
Claim under test = 'SEND_ZC on the real fabric is a win, so keeping the default ON is right'. Expected: in every one of the 3 amd64 pairs the ON arm shows lower cpu_per_GB on ws-large-echo by >= 5 % (e.g. mean_cpu_pct drops from the ~96 % plateau of the v1.5.8 data toward epoll's 53-58 % at the same 27-29.7 K msg/s, or rps rises above ~29.7 K at the same CPU); bpftrace shows opcode-47 submissions in the thousands per second during the cell with skb_copy_ubufs ~0 (true zero-copy on the bond); the controls ws-echo and get-json differ between arms by <= 1.6 % rps and <= 0.6 CPU points. On arm64 ON rps >= OFF rps + 3 %.
Expected if the claim is FALSE
Two distinct falsifications. LOSS: the OFF arm has lower cpu_per_GB by >= 5 % and/or higher rps by >= 3 % in 3/3 pairs — most plausibly OFF brings the io_uring column's ~96 % CPU down toward epoll's ~55-58 % at equal throughput (the NOTIF-per-send + buffer-hold cost, or a copy-fallback where bpftrace shows skb_copy_ubufs count ~= opcode-47 count, meaning the kernel copies anyway and ZC only adds a second CQE). NOISE: |Delta cpu_per_GB| < 5 % and |Delta rps| < 3 % in every pair with inconsistent sign, controls moving as much as the treatment; then the default is immaterial for the published grid and #465's residual question is closed as 'no measurable effect on this fabric', which is also a real result. Either way the ~1.7x CPU-per-byte gap between io_uring and epoll on this cell becomes attributable (to ZC, or explicitly NOT to ZC -> a separate issue on the detached-write funnel).
Negative controls
Three built into every dispatch: (1) ws-echo — same detached inline-write/detachQueue path, 256 B frames < 4096 so useSendZC is false (send_zc_gate_test.go:24-25 pins this) — must show no arm difference; (2) get-json — linked SEND only, ZC unreachable — must show no arm difference; if either control moves as much as ws-large-echo the delta is run-level drift, not ZC. (3) Arm verification: server.log must read send_zc=false in the OFF arm; the bpftrace opcode-47 count must be 0 in the OFF arm and > 0 in the ON arm (proves the env reached a root-launched SUT and that the cell actually exercises the branch — without this a 'noise' verdict could just mean the toggle never applied or the branch never fired). Existing-data sanity anchor: the ON arm at profile=fast should reproduce the v1.5.8 fast-run band (28.8-29.7 K msg/s, 96.3-96.8 % CPU) within the v1.6.0 engine changes; a large departure from that band in BOTH arms indicates something else moved (#572/#574/#525 etc.), to be reported, not folded into the ZC verdict.
Cost
Per dispatch with the cells passthrough (3 cells x 1 column x 1 arch): setup runner ~5-10 min + mage Deploy (Go build of servers/celeris only) ~5-10 min + 3 cells x ~2 min (budget model 10+35+5+12 = 62 s/cell, ~2x observed under-count) + merge/upload/cleanup/teardown ~10-15 min = ~35-45 min, i.e. ~4-4.5 h cluster wall-clock for the six amd64 dispatches; the same again for arm64 (msr1) = ~9 h total for both arches, versus ~74 h for the release bench. WITHOUT the cells passthrough (column-scoped only via the existing competitors input) each dispatch runs the column's ~29 capability-gated scenarios (~1 h at fast + overhead) = ~1.5 h/dispatch, ~9 h per arch, ~18 h both — the controls then come for free. Engineering: ~1-2 h for the passthroughs + a unit test. Statistical cost: three pairs give a 3/3 sign requirement (one-sided p = 0.125 on sign alone) — the decision rests on the magnitude thresholds being 3x the measured same-window spread (rps <= 1.6 %, CPU <= 0.6 pt over five byte-identical v1.5.8 runs), not on the sign test.
Not measurable, and why
(1) Whether the NIC/bond actually zero-copies cannot be observed from celeris: production SQEs never set IORING_SEND_ZC_REPORT_USAGE (sqe.go:175-186 sets only opcode/fd/buf), so NOTIF CQEs carry no ZC_COPIED bit; only the out-of-harness kernel probe (skb_copy_ubufs / tracepoint opcode filter, root on the SUT) can tell copy-fallback from true ZC, and the harness cannot record it. (2) The fraction of ws-large-echo bytes that go through SEND_ZC vs the inline unix.Write fast path (worker.go:1567) is not counted anywhere — EngineMetrics (engine.go:335-351) has bytesWritten but no ZC counter — so a NOISE verdict could mean 'ZC rarely fires' rather than 'ZC is free'; the bpftrace step or a small celeris counter is the only way to separate these. (3) HTTP/1.1 and h2 large responses are structurally ZC-unreachable in the bench (linked SEND, WRITEV bodies, inline h2 writes), so this A/B says NOTHING about ZC for HTTP; the only configuration where every >= 4 KiB send would be ZC is multishot recv (CELERIS_IOURING_MULTISHOT_RECV=1, bufRing != nil -> flushSend unlinked), which the bench does not run — a second A/B dimension deliberately out of scope. (4) Fixed-file mode (off by default, celeris#541/#553) and the adaptive columns' post-promotion behaviour are not covered. (5) The CI memlock-limited ENOMEM fallback path (worker.go:2434, 2521; 8 MB memlock on GitHub runners) is not exercised on the cluster (memlock unlimited). (6) Cross-dispatch rps is confounded by window length (35 s vs 90 s) and by any parallel arch pass, so only within-pair comparisons at identical profile/target are valid; a single ON run compared against the historical v1.5.8 numbers is NOT a measurement. (7) mean_cpu_pct averages the whole cell-guard window including idle edges (dilutes both arms equally) and is host-wide, so it cannot attribute CPU to the SUT process vs kernel softirq — a
perfsplit is out of harness.Code changes needed
Required (probatorium, harness only; no celeris change needed to run the A/B): (1) SUT env passthrough — benchmark-tier.yml: add workflow_dispatch input
sut_env(KEY=VALUE, default "") -> job env BENCH_SUT_ENV; mage_bench.go Bench() buildArgs (mage_bench.go:451-483): when BENCH_SUT_ENV is set, append --extra-vars with a JSON dict {"bench_sut_env": {KEY: VALUE}}; ansible/tasks/run_bench_cell.yml 'start competitor server in background' (lines 199-272): turn the hard-codedenvironment:mapping (207-235) intoenvironment: "{{ sut_base_env | combine(bench_sut_env | default({})) }}"and also log the merged env into server.log (or echo it in the launch shell) so the arm is recorded next to the send_zc= line. (2) Cell-scope passthrough — benchmark-tier.yml: inputcells(default "") -> env BENCH_CELLS; mage_tier.go setBenchEnvFromProfile (lines 218-244): honour a pre-set BENCH_CELLS the same way it already honours BENCH_DURATION/BENCH_WARMUP instead of unconditionally os.Setenv("BENCH_CELLS", budget.CellsGlob(p)); FitWithin still projects p.Cells=813 so it passes trivially — acceptable for a scoped run, note it in the log. Unit tests: a table test for the merged extra-vars and one for the BENCH_CELLS precedence. Optional but recommended for interpretability: (3) celeris — two monotonic counters on Worker (zcSendsSubmitted in prepSendSQE's ZC branch worker.go:3973-3977, zcNotifs in handleSend worker.go:2370) surfaced through EngineMetrics (engine.go:335-351) and thus /debug/vars, which the bench observer already scrapes for celeris cells (run_bench_cell.yml:421) — this makes the 'did ZC actually fire' question harness-recordable instead of bpftrace-only; small, additive, covered by extending send_zc_gate_test.go. (4) Optional: record bench_sut_env in the merged Document's environment (mage_bench.go:1788-1806) so an arm is identifiable from results.json alone. If the verdict is LOSS: celeris probe.go:258-259 'auto' -> return false (default OFF; explicit on still honoured), doc.go:9-13 wording, and a repin — that is the follow-up, not part of the measurement.Corrections from the adversarial review (binding)
Review 1: refuted
Why: Two independent defects each make the proposed A/B unable to separate "SEND_ZC is a win/loss" from "SEND_ZC never mattered", and the design's decision rule then converts the guaranteed NOISE outcome into a verdict.
The decision metric does not exist per cell.
resources.summary.mean_cpu_pctis the mean of ONE mpstat run over the WHOLE competitor column: in the cluster harness a "cell" is(run_index, competitor_slug)(probatorium ansible/tasks/run_bench_cell.yml:2-6), the server is launched once andmpstat -P ALL 1 {{cell_guard_seconds}}runs once per column (lines 384-389), andaggregatePerCellResultsstamps that single aggregate onto every scenario the runner expanded ("The same aggregate applies to every scenario ... since the observer scopes to the whole cell process", mage_bench.go:996-1004). Verified in the published data the design cites: in docs/results/v1.5.8/20260805/x86_64/summary.json every celeris column has 27 scenarios, exactly 1 distinct resources summary, byte-identical series across ws-large-echo and get-json, spanning 1334 s (8024 s in the 20260829 full run). So the motivating "existing signal" — io_uring 95.6-96.8 % vs epoll 53-58 % "on precisely the ZC-eligible cell" — is the column-wide average over churn-close, driver-*, get-json, etc., and says nothing about ws-large-echo. Under the proposed 3-cell scoping the treatment cell and BOTH controls would share one CPU scalar: the control test ("ws-echo/get-json must move <= 0.6 CPU pt") is unevaluable, and any treatment delta is diluted by two control cells plus cooldown/idle gaps. cpu_per_GB as specified cannot discriminate ON from OFF. rps cannot rescue it on amd64: at 29.7 K msg/s x 64 KiB x 2 directions the cell is near the 2x10G bond's line rate (std reaches 35 K), which is why the design itself leans on CPU.Under the bench configuration the ZC branch almost never carries bytes, so the toggle is nearly inert. Detached WS egress first tries an inline
unix.Write(cs.fd, cs.writeBuf)on the dispatch goroutine and only hands the PARTIAL remainder to the ring (celeris engine/iouring/worker.go:1567-1583 guarded closure; ring path = drainDetachQueue -> flushSend -> prepSendSQE -> useSendZC, worker.go:3896-3991). loadgen's ws-large-echo is strictly ping-pong —exchangewrites one frame then blocks until its echo arrives (loadgen ws.go:138-163) — so each conn has at most one 64 KiB frame in flight and an empty send queue when the echo is written; the SUT runs kernel-default autotuned tcp_wmem (bench_tuning.yml touches no wmem/rmem sysctl). In steady state the whole frame goes inline; SEND_ZC only fires for the rare short-write remainder that is also >= 4096 B. The inline path landed in 0af8831 (2026-07-03) and is inside v1.5.7 and v1.5.8, so even the cited 96 % number was produced with ZC essentially unreachable. The design concedes this in not_measurable(2) but its decision rule maps NOISE to "leave the default ON, close io_uring: probeSendZC can never detect copy-fallback — SEND_ZC enabled on every host regardless of NIC #465 as no measurable effect" — the result the run would produce whether a ZC send costs 0 % or 50 % more CPU than a SEND, because there is no ZC-fired count to distinguish "ZC is free" from "ZC never executed". The bpftrace count is optional and one-arm-only, and the celeris ZC counters are optional.Minor: neither the
sut_envnor thecellsworkflow input exists yet, so step 3's command line cannot run today (the design acknowledges this). The arm-identity check is sound:profile.SendZC = enabledprecedesSelectTier(profile)(engine.go:97-103), so the loggedsend_zc=is the effective post-env value.Change to the design: Make both legs discriminating before spending cluster time:
A. Per-cell CPU. Never use the summary scalar. Compute per-scenario CPU by windowing the raw 1 Hz
cpu.log(wall-clock rows; report/resources.go ParseMPStat already returns the series) against the runner's per-scenariostarted_at/completed_atrecorded in results.json (cmd/runner/main.go:1093-1096), trimming the warmup window; or have the harness restart mpstat per scenario. Add SUT-process CPU (utime+stime from /proc//stat, the observer already owns the PID: run_bench_cell.yml:420) and mpstat's %soft so a ZC cost (kernel NOTIF/skb_copy_ubufs work) is attributable rather than host-wide. Define cpu_per_GB and the control comparisons on these windowed numbers only. Before dispatching anything, apply the same windowing to a retained v1.5.8/20260829 run artefact to learn whether io_uring's per-cell CPU on ws-large-echo actually differs from epoll's — the current 96 % vs 55 % motivation is a column average and must be retracted.B. Make ZC carry the bytes, and prove it. (1) Promote the celeris counters from optional to REQUIRED: zcSendsSubmitted, zcNotifs, inlineBytes (the unix.Write fast path) and ringBytes (flushSend/WRITEV) on Worker, surfaced via EngineMetrics -> /debug/vars, scraped by the observer per cell; a verdict is only valid if ringBytes/(inlineBytes+ringBytes) >= ~50 % and zcSendsSubmitted is in the thousands/s in the ON arm and 0 in the OFF arm; otherwise the run has answered a different question. (2) Because the production ping-pong path takes the inline fast path, run a 2x2: {SEND_ZC on/off} x {inline egress on/off}. Add a celeris env knob (e.g. CELERIS_IOURING_INLINE_EGRESS=off) that skips the unix.Write fast path so every detached >= 4 KiB send reaches prepSendSQE; the inline-off pair is the only one that measures SEND vs SEND_ZC on the fabric, and the inline-on pair measures whether the default matters in production (expected: it does not fire — which is then a real, counted result, not an inferred one). Alternatively (no celeris knob) add a large-frame fan-out cell (ws-hub-broadcast with 64 KiB payload, the shape celeris's own broadcast_egress_correctness_linux_test.go uses) where per-conn queues fill and the ring path genuinely carries the data. (3) Keep bpftrace opcode-47/26 + kprobe:skb_copy_ubufs, but run it in BOTH arms of at least one pair and make it a precondition (opcode-47 > 0 ON, == 0 OFF), plus
ethtool -k bond0 and slaves | grep scatter-gather, so copy-fallback vs true ZC is on record.C. Keep ABBAAB ordering, one arch per dispatch, server.log
send_zc=arm check, and the harness passthroughs (sut_env, cells). Restate the decision rule on the windowed per-cell metrics and the ZC-fired precondition: WIN/LOSS thresholds as before; NOISE is only a verdict when the ON arm demonstrably routed >= 50 % of ws-large-echo bytes through SEND_ZC; a NOISE result with a low ZC fraction closes #465's production question ("the default is immaterial because the branch does not fire on the echo path") but leaves the fabric question open for the inline-off pair.Review 2: refuted
Why: The design's DECISION METRIC (cpu_per_GB = mean_cpu_pct / bytes on ws-large-echo) and the premise that motivates the whole A/B ("io_uring at 95.6-96.8 % SUT CPU on precisely the ZC-eligible cell vs epoll 53-58 %") both rest on
resources.summary.mean_cpu_pctbeing a per-(column, scenario) number. It is not. In the harness a "cell" is one COLUMN's whole runner pass: bench.yml:266-282 builds the schedule as [(0, competitor)...], run_bench_cell.yml:42 sets cell_dir =<run_dir>/00-<competitor>, the mpstat sampler (run_bench_cell.yml:384-389) and the observer (:404-426) are started ONCE per that dir for cell_guard_seconds, and the single runner invocation (:514-525) runs every scenario matched by-cellsinside it. The merge then copies the one aggregate onto every scenario — mage_bench.go:996-1004, in its own words: "The same aggregate applies to every scenario the runner expanded in this cell, since the observer scopes to the whole cell process." The published data proves it: in all six v1.5.8 summary.json files every column has exactly ONE distinct mean_cpu_pct across its 27-29 scenarios (e.g. celeris-iouring-h1-async 20260829 amd64: churn-close, get-json, ws-echo, ws-large-echo ... all 96.283 %; epoll-h1-async all 53.31 %), and the "ws-large-echo" series spans 1334 s (fast) / 8024-8850 s (90 s runs) — the whole column — not one ~35-90 s cell. So the "1.7x CPU per byte on the ZC-eligible cell" is an artefact of reading a column-wide mean (dominated by 27 other scenarios incl. get-json at 1 M rps and churn-close) as a per-cell value; the io_uring-vs-epoll CPU gap is column-wide (consistent with worker.go:1252 adaptiveTimeout()==0 whenever dirtyHead != nil, i.e. a busy loop under any pending send) and says nothing specific about SEND_ZC. Consequences for the proposed measurement as written: (a) with the plannedcells='ws-large-echo/*,ws-echo/*,get-json/*'passthrough, all three scenarios run in ONE column pass and share ONE cpu.log, so the treatment cell and both "negative controls" would report the IDENTICAL mean_cpu_pct by construction — the CPU control is vacuous and cannot flag drift; (b) that shared number would be the average over get-json (CPU-saturated at ~1 M rps regardless of ZC), ws-echo, ws-large-echo and the idle/warmup edges, so a real ZC-driven CPU change on ws-large-echo (which contributes ~1/3 of the window) is diluted at least 3x before the 5 % threshold is applied, biasing the rule toward NOISE — and the design explicitly closes #465 on NOISE as "no measurable effect on this fabric"; (c) the 5 %/3 % thresholds were calibrated against "same-window spread ... CPU <= 0.6 pt over five v1.5.8 runs", which is the spread of a column-wide mean over ~27 scenarios, not the per-cell noise floor, so the threshold itself is mis-sized. Secondary, non-load-bearing errors: the "expected_if_true" anchor "mean_cpu_pct drops from the ~96 % plateau toward epoll's 53-58 %" compares column-wide means; andresources.seriescpu_pct points are the same column-wide 60-point downsample. What is NOT refuted: saturation_mode_rps on ws-large-echo is a genuine per-scenario observable (loadgen ws.go:138-153 is a synchronous write-then-read per conn, so msg/s tracks SUT turnaround; std reaching 34.8-35.4 K on the same fabric shows io_uring's 26-29.7 K is not wire-capped), the send_zc=true|false log-line arm check, the code-path analysis (inline unix.Write fast path at worker.go:1567 with ring SEND -> useSendZC(n>=4096) only on the EAGAIN/partial remainder; flushSendLink 'never ZC' for H1), and the ABBAAB ordering. The design would produce numbers, but its CPU-per-byte verdict would be about the column, not the cell.Change to the design: Make the CPU observable per-scenario before running anything, and re-base the decision rule on it. Concretely: (1) In the merge (mage_bench.go readCellResources / report.SummarizeResources) window the column's cpu.log and observer.sqlite per scenario using the per-scenario started_at/finished_at the runner already writes (cmd/runner/main.go:304/339, report/schema.go:296-297; mpstat rows carry wall-clock times, observer rows carry ts_unix), and attach the sliced ResourceStats to each scenario record instead of the shared aggregate; alternatively, and with zero merge changes, dispatch ONE scenario per column pass (
cells='ws-large-echo/*'only, and separate dispatches for the ws-echo and get-json controls) so the column-wide mean IS the cell mean — but then compute the mean over the measurement window only (exclude the warmup/idle edges via the timeseriestin run0//*.json), because the guard window is ~5/4 x budget + 300 s and mostly idle for a 35 s cell. (2) Replace host-wide mpstat as the primary with per-process SUT CPU (utime+stime of the server.pid from /proc, which the observer already sidecars, orpidstat -p) so worker busy-spin/softirq is at least separable; keep mpstat 'all' as secondary. (3) Re-derive the noise floor and the 5 %/3 % thresholds from per-cell repeats (>=3 identical-arm ws-large-echo runs), not from the column-wide spread. (4) Because the io_uring worker spins whenever dirtyHead != nil (adaptiveTimeout()==0), report and reason on rps as the primary decision metric on both arches (it is per-scenario and not wire-capped), with per-scenario CPU as the secondary; do not close #465 on a CPU NOISE verdict unless the per-scenario CPU is shown to move with a known perturbation (e.g. the ws-echo vs ws-large-echo difference) so the metric is proven sensitive. (5) Make the ZC-fired counter (zcSendsSubmitted / zcNotifs via EngineMetrics -> /debug/vars, scraped by the observer) REQUIRED rather than optional, so NOISE can be distinguished from "ZC never fired" on this path; the bpftrace step can stay optional. (6) Correct the premise in the issue write-up: the 96 % vs 55 % io_uring-vs-epoll CPU gap is column-wide across all ~27 scenarios and belongs to a separate io_uring loop-CPU issue, not to #465.