You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Measure: capture enough at the next h2c_hang / ws_handshake_fail to root-cause it from the artifact, and try to force recurrence #588
Filed from the release-readiness sweep of 2026-09-13 under the standing rule that a closure counts only when it is MEASURED. Design produced by a read-only design pass and corrected by two adversarial reviews; the corrections are binding and listed last.
Closure under test
v1.5.11 24h soak (probatorium run 34003151551, celeris v1.5.11 pins, probatorium f8e5d3c): auth_session_ratelimit/iouring/arm64 tier_1.h2c_churn.h2c_hang=1 (cause=timeout, max_elapsed=10000 ms) and auth_session_ratelimit/epoll/arm64 tier_1.ws_torture.ws_handshake_fail=1 (pre-cause-split); neither root-caused; 0 recurrences in the 96 cells of soaks 34368601903 and 34616620237.
What the closure claims
NOT closable as root-caused, and NOT closable as "gone". The two original events are permanently unattributable: the artifact kept only the counters (verified: the msr1 cell-03 directory in the downloaded v1.5.11 artifact held nothing but incidents/20260906-043231-I-MEM-1; no tally sidecar, no error string, no timestamp, no engine log). 0-in-96 later cells is statistically compatible with the observed per-cell rate (~1 h2c hang per ~120k reads, ~1 WS failure per ~47.6k handshakes at concurrency 50 — the memory's own "0-in-55k and 1-in-55k are the same rate" lesson applies). What CAN be closed by measurement is the apparatus claim "the next occurrence will be root-causable from the artifact alone" — that claim is FALSE today, and the measurement below turns it TRUE and verifies it with a fault-injected negative control before the next soak. Separately, a directed 6 h cell can force the rate into a measurable regime (~6 events per cell per counter if the observed rate holds).
What was argued rather than measured
"The soak's single h2c_hang was a plain responsiveness stall" (memory absolute-gate-v159-release.md, posted to probatorium#298): argued from config only — every refapp pins Protocol: celeris.HTTP1 (validation/refapp/auth_session_ratelimit/main.go:185) and config.go:289-294 infers EnableH2Upgrade=false, so the preamble is a plain GET / (h2c.go:337-358) that hits session+ratelimit (main.go:215-245; / is not routed, so 404 or 429 = "declined"). Correct that it is not an upgrade path, but stall-vs-deaf-listener-vs-server-header-timeout was never discriminated: h2c.go:185-194 says in so many words that at the then-10 s read budget "it is not recoverable from the artifact whether the server stalled or merely hit its own header deadline" (ReadHeaderTimeout default 10 s, config.go:92-102; enforced in iouring worker.go:1479/2796 and epoll loop.go:297). The budget was raised to 20 s in feat: per-handler async dispatch (sync/async per route + group, all engines × all protocols) #302 (merged 2026-09-07, AFTER this soak).
The WS failure's cause class is unknown: the eof/timeout/reset/status/other split (ws.go:56-72, 132-136) also landed in feat: per-handler async dispatch (sync/async per route + group, all engines × all protocols) #302 after the soak. By construction it cannot be a 401/429 (/ws skips both middlewares, main.go:214-245) and the walker's request is well-formed and Origin-less over http, so websocket.go:64-101 yields 101; a non-101 would have to be a 400/403/426 from a corrupted request — nothing in the artifact can say.
The strongest prior is celeris#493 (closed v1.5.10): the v1.5.9 soak's 6 h2c hangs + 40 WS + 8 SSE handshake failures on the SAME cell type/arch were a once-a-minute worker-event-loop freeze (store sweep under shard lock with inline dispatch), proven by goroutine dumps inside a stall. The bounded sweeps shipped (middleware/store/memory.go:300-310 sweepBatch=2048; middleware/ratelimit/shard.go:164-213 with runtime.Gosched). v1.5.11 still recorded 1+1 on that cell type — residual sub-threshold stall or a different mechanism; no evidence either way. celeris#284 (v1.4.8) is the same symptom class never reproduced in isolation, whose "suggested next step" — log per-attempt elapsed and failure mode — is still not implemented.
What the harness discards today (all read): h2c.go:286-295 calls recordHang(err, elapsed) which keeps only a cause class and a GLOBAL max elapsed (h2c.go:204-226); the mode, walker seed, err.Error(), conn.LocalAddr(), wall-clock instant and per-event dial/write/read durations are dropped; dial failures (h2c.go:242-246) and write failures (252-258) return with NO counter. ws.go: fireWSTorture runs dial+write+status+headers under ONE 2 s deadline (wsMaxHold, ws.go:192, 233-241) versus h2c's 20 s — so a 2-10 s server stall fails WS but not h2c; the non-101 status text is counted (ws.go:272) but not kept; dial failures dropped (236-239). Neither counter is in runner.go's reactive fire() list (runner.go:953-975 covers only adv_wrong_accepted, h2c_crashed, ws_accepted_bad_frame, ws_hang_no_close, liveness, hang), so no incident dossier — no goroutine.pprof, /proc//stack, fd list, dmesg (forensics.go:31-114) — is ever taken at the instant. The refapp's stdout+stderr goes through remote/local.go:63-64 into superviseStderr's 80-line ring (liveness.go:376-380) which is attached ONLY on process death (liveness.go:457-475); the engine Warn/Error lines that would name a mechanism — epoll "accepted fd exceeds conn table cap; dropping" (loop.go:770), EMFILE/ENFILE (loop.go:740-748), io_uring "re-create listen socket" (worker.go:653) — are discarded for a live refapp (the refapp sets no Config.Logger, main.go:181-191, so they go to slog.Default → stderr). EngineMetrics.ErrorCount (engine/engine.go) is bumped on those paths (loop.go:747/753/766; worker.go:1293/1333) but NOT exported by debugvars (debugvars.go:262-290 exports only active_conns/adaptive_switches/engine) and not in the 1 Hz series (series.go:21-25: ts, goroutines, heap*, rss, accepted, closed, active). The io_uring SQ-full accept drop (worker.go:3586-3591) sets acceptRearmPending without bumping any counter. The cell document has no per-cell start/ready time (ValidationCellResult in report/schema.go; StartedAt is run-level, matrix.go:493), so the io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493 "stall at the refapp's start second every minute" phase signature cannot be tested even if a timestamp existed.
The gate (report/gate.go:303-326, 452-463) fails on h2c_hang and ws_handshake_fail totals and folds the cause counters into the message (causeSuffix) — correct behaviour, kept.
Measurement
Observable
Two things, in order. (A) Attributability: for every h2c_hang and ws_handshake_fail event the cell document must carry a per-event record (bounded ring of 16, modelled on #327's sse_early_errs at sse.go:73-79,107-145 / schema.go:618-623): UTC ts (RFC3339Nano), ms since refapp ready, walker seed+id, churn/torture mode, cause class, verbatim err.Error(), non-101 status line (WS), walker conn.LocalAddr() (the engine<->client join key from celeris#562), THIS event's dial/write/read durations and bytes read, plus validator-side GC pause since last fire (runtime/metrics) — and, alongside it, an incident dossier taken within 2 s (goroutine.pprof, /proc//stack, fd list, dmesg), the refapp stderr ring flushed to /refapp_stderr_tail.txt, celeris.engine_error_count added to /debug/vars and the 1 Hz series, and ready_at in the cell doc. (B) Rate/mechanism: count of events per 6 h directed cell per counter, their timestamps, and the series columns accepted/active/engine_error_count in the +/-10 s window around each event.
Where it runs
(A1) Unit, probatorium validation/ under -race on Linux: extend newFakeH2CServer (h2c_test.go:89-113) and newFakeWSServer (ws_test.go:163-217) with policies {silent 25 s, accept-then-close, RST, slow-101 at 3 s} and assert the per-event record fields. (A2) Apparatus negative control in a Linux container with --security-opt seccomp=unconfined --ulimit memlock=-1 (the imem1 rig at scratchpad/imem1/{drive.sh,run.sh}): one auth_session_ratelimit io_uring cell, concurrency 50, 20 min; at a known instant T a sidecar sends SIGSTOP to the refapp for 3 s (and once for 25 s), then SIGCONT. (B) Probatorium cluster weekend tier, scoped: VALIDATE_MATRIX=1 VALIDATE_MATRIX_REFAPPS=auth_session_ratelimit VALIDATE_MATRIX_ENGINES=iouring,epoll,std SOAK_DURATION=18h VALIDATE_TARGET=both VALIDATE_CONCURRENCY=50 -> cellBudget = 18h/3 = 6 h per cell (cmd/validator/matrix.go:403), both hosts in parallel (VALIDATE_PARALLEL=1). The dispatch currently exposes only duration/resume_from (matrix-weekend-tier.yml:15-30), so this needs either a workflow input or a mage Soak from the runner host; VALIDATE_GATE_EXPECT_CELLS must be 6 (not 48) for that run.
Procedure
Land the code changes (see code_changes_needed) with the unit tests; run go test -race ./validation/... ./report/... on Linux. 2. Run the fault-injected container cell: verify (a) the 3 s SIGSTOP produces ws_handshake_fail_timeout events whose ts is within [T, T+3 s], h2c_hang=0 (3 s < 20 s), sse early closes/timeouts at the same T, series accepted flat for 3 rows, a dossier dir incidents/-I-WS-HANDSHAKE with goroutine.pprof, refapp_stderr_tail.txt present; (b) the 25 s SIGSTOP additionally produces h2c_hang_timeout with per-event read elapsed ~20000 ms and I-H2C-HANG dossier; (c) with no injection over 20 min all new rings are empty and the gate still PASSes (guard test like deps: bump the all-go-deps group across 3 directories with 5 updates #327's fails if the rings are dropped from schema or gate). 3. Dispatch the scoped 18 h soak (6 h x 3 engines per host). 4. From the artifact: tally events per cell per counter; for each event join ts against the series and against the SSE/adversarial rings; compute (ts - ready_at) mod 60 s; read the dossier's goroutine.pprof for the worker goroutines' state; read refapp_stderr_tail for engine Warn/Error lines. 5. Compare iouring vs epoll vs std per host.
Expected if the claim is TRUE
TRUE = a server-side mechanism exists at the observed rate. In each 6 h iouring/epoll cell expect ~6 h2c hangs (5 walkers x 10/s x 21600 s x 2/3 non-RST reads = ~720k reads at ~1/120k) and ~6 WS failures (2 walkers x 6.67/s x 21600 s = ~288k at ~1/47.6k) with 95% bands roughly [2,12]; events CLUSTER in time across slices (h2c, WS, SSE early-close, adv_hang_until_timeout within the same few seconds); h2c per-event read elapsed ~20000 ms with cause=timeout and WS ~2000 ms timeout for a stall, OR small-elapsed eof/reset for a drop; series accepted flat (or active spiking) across the event window; goroutine.pprof shows worker goroutines off the io_uring_enter/epoll_wait syscall (blocked on a lock or running a sweep) for the stall class, or the listen worker idle with no accept SQE for the deaf-listener class; celeris.engine_error_count steps at the event for the fd-cap/EMFILE class; if #493-residual, (ts - ready_at) mod 60 clusters within a few seconds of one phase; std shows 0 (no worker event loop). The error string discriminates further: 'i/o timeout' = never answered; 'EOF' at ~10000 ms = server ReadHeaderTimeout closed an accepted-but-unread conn; 'connection reset by peer' small elapsed = fd dropped at accept.
Expected if the claim is FALSE
FALSE = apparatus/transport artefact (the ~1/9000 dial-loss class the SSE analysis already found on all engines). Events are uniformly spread, do not coincide with SSE/adv events, accepted never flattens, goroutine.pprof is unremarkable, engine_error_count does not move, the phase histogram is uniform, the new h2c/ws dial_fail and write_fail counters show a comparable rate on std, and the error strings are client-side (dial i/o timeout, write: broken pipe) with validator GC pause > 0 at the instant. Also possible: 0 events in 6 h x 3 cells per host (~2.2M h2c reads, ~860k WS handshakes total) which bounds the rate below ~1/1M and says the v1.5.11 pair was a rate that #574/#572-era changes removed — that is a measured statement, unlike today's 0-in-96.
Negative controls
Three: (1) the fault-injected SIGSTOP cell above (known instant, known duration) must produce exactly the predicted counters/records/dossier with timestamps within 2 s of T — this proves the capture path works before it is relied upon; without injection the same cell must produce empty rings and a PASS (proves no false positives; run it also with the rings assertion inverted to confirm the guard test fails when the fields are removed, per RULE: always run the negative control); (2) the std-engine 6 h cell in the same dispatch: std has no worker event loop and no fd-cap table, so a server-side stall/deaf-listener mechanism must NOT appear there while an apparatus artefact appears at the same rate; (3) unit fake-server policies assert the record content for each cause class so a future refactor cannot silently drop the error string again (the #327 guard-test pattern).
Cost
Unit tests: ~1-2 min per package under -race. Container fault-injection: ~30 min wall (2 x 20 min cells) on any Linux box with seccomp=unconfined. Directed cluster run: 18 h wall on both hosts in parallel (6 cells total, 3 per host), exclusive cluster time; if 0 events, one repeat at VALIDATE_CONCURRENCY=100 (h2c 10 walkers, WS 5) for another 18 h before bounding the rate — note that changes the load profile so its I-MEM verdicts are not comparable to the standard soak. msr1 is the runner host: nothing may touch it during the run. Code: ~1 day incl. schema bump and tests.
Not measurable, and why
(1) The two original v1.5.11 events: no error string, timestamp, stderr tail or dossier exists anywhere (artifact verified; /tmp/soak3 is empty; the run artifact 340 KB is downloadable but contains only the run document and I-MEM-1 dossiers per the memory's own audit). They stay unattributed forever. (2) Stalls shorter than the 2 s tallyCB interval (runner.go:1001) will have a dossier taken after the stall ended; only the per-event record + 1 Hz series catch them, and a <1 s stall cannot fail either walker anyway (WS budget 2 s). (3) Kernel accept-queue occupancy at the instant (ss -ltn / /proc/net/tcp) is not captured by forensics.go; 'deaf listener' vs 'worker stalled after accept' is inferred from elapsed+EOF patterns and goroutine state, not observed directly — adding an ss -ltnp snapshot to the dossier would close that but is a further change. (4) The validator shares the host and loopback with the refapp, so a validator-process scheduler stall is indistinguishable from a server stall from the walker's clock alone; the record's validator GC-pause field and the cross-slice clustering bound this but cannot exclude a host-wide stall (both processes freeze together). (5) A rate below ~1/1M per counter is not resolvable in one 6 h cell; only repeated directed runs can push the bound down.
Code changes needed
probatorium (all under validation/ unless noted): (1) h2c.go: recordHang(err, elapsed) -> recordHang(ev h2cEvent) with a mutex-guarded ring of 16 {ts, since_ready_ms, walker, mode, cause, err, local_addr, dial_ms, write_ms, read_ms, n_read, validator_gc_pause_ns}; count dial/write failures as h2c_dial_fail / h2c_write_fail (informational, not gated); snapshot() emits h2c_hang_events. (2) ws.go: same for recordHandshakeFail (+ status line text), ws_dial_fail / ws_write_fail; keep the 2 s wsMaxHold but split the dial deadline from the handshake read deadline so the record can say which leg expired. (3) runner.go Tier1Summary projection + report/schema.go: new H2CHangEvents / WSHandshakeFailEvents []struct fields next to SSEEarlyErrs (schema.go:618-623), schema minor bump; report/gate.go causeSuffix prints the first event's err/elapsed (informational). (4) runner.go tallyCB: fire(properties.IH2CHang.ID, ..., snap.H2CChurn.Hang > 0) and fire(properties.IWSHandshake.ID, ..., snap.WSTorture.HandshakeFail > 0) with RecordOnly=true so the cell runs on but a dossier is written; add the two Specs to properties/i_tier1_walker.go and registry.go. (5) liveness.go: expose the stderr ring via a tailSnapshot() on livenessTally; runner.go writeIncidentDossier writes refapp_stderr_tail.txt into the dossier and tier1 end-of-cell writes /refapp_stderr_tail.txt unconditionally (80 lines, bounded); optionally tee the refapp's merged output to /refapp.log with a size cap in remote/local.go. (6) refapp/internal/debugvars: export celeris.engine_error_count from EngineMetrics.ErrorCount; checker/poll.go Snapshot field; series.go add the column; record the ready-banner time as ready_at in the cell doc (runner.go/matrix.go). (7) .github/workflows/matrix-weekend-tier.yml: dispatch inputs refapps/engines mapped to VALIDATE_MATRIX_REFAPPS/ENGINES and VALIDATE_GATE_EXPECT_CELLS derived from the selection. (8) Tests: h2c_class_test.go / ws_test.go per-event assertions for silent/close/RST/slow policies; a gate guard test that fails if the event rings vanish from schema or the gated totals stop being gated. celeris (optional but raises signal): bump errCount when prepareAccept hits a full SQ ring (engine/iouring/worker.go:3586-3591 currently only sets acceptRearmPending) and expose fdCapDrops (engine/epoll/loop.go:129-135) through EngineMetrics.ErrorCount, so the exported counter moves on exactly the suspected deaf-listener paths.
Corrections from the adversarial review (binding)
Review 1: refuted
Why: Part (B) cannot discriminate its two hypotheses at the sample size proposed. The whole basis is ONE h2c event and ONE WS event; a single Poisson observation has a 95% rate interval of roughly [0.025, 5.6] x the point estimate, so the 'expected_if_true' band '[2,12] events per 6 h cell' is fiction — the honest expectation for a 6 h cell (720k h2c reads / 288k WS handshakes) is anywhere from 0.05 to ~11 events, and the pooled data already in hand (1 in 3 same-config cells, i.e. ~1/360k h2c and ~1/143k WS; or 1 in 5.8M h2c / 1 in ~1M WS if uniform across cells) gives lambda ~2 with P(0)=0.14, or lambda 0.04-0.3 with P(0)>=0.74. A 0-event result is therefore compatible with 'mechanism still present', 'mechanism removed' and 'apparatus artefact' alike, and the design explicitly plans to read 0 events as the FALSE branch ('bounds the rate below ~1/1M ... says the v1.5.11 pair was removed') when it bounds nothing the 96 routine cells have not already bounded. The std 'negative control' is powerless for the same reason: at lambda 0.04-0.3 std shows 0 under BOTH hypotheses (P>=0.74), so 'must NOT appear on std' is satisfied by chance. It is also mis-specified: the h2c GET / goes through session+ratelimit on std too (main.go:214-245 skips only /ws,/events), so a lock-hold stall WOULD show on std for h2c and not for WS. Part (A)'s key observable does not survive contact with the capture path: /debug/vars and /debug/pprof are served THROUGH the engine under test (debugvars.go:217 srv.Pre(v.Handler(), pprof.New())), so for the process-wide worker-freeze class the design says expected_if_true hinges on ('goroutine.pprof shows workers off the io_uring_enter/epoll_wait syscall'), the 500 ms poll and the 5 s pprof fetch both fail during the stall and the dossier holds goroutine.pprof.missing; celeris#493 was only proven because the diag build had a pprof SIDE-listener and a 200 ms ss sampler — neither is in this design (ss is explicitly deferred as 'a further change'). Even with a side listener, alertedCounters fires the dossier once per counter per cell and only after the walker's budget expired (>=2 s WS, >=20 s h2c, +<=2 s tallyCB), so for the observed 1.5-8 s stall class the h2c dossier is always post-stall and the WS one is post-stall for anything under ~4 s: the record says 'timeout', the dossier shows a healthy process, and 'server stall' vs 'host-wide stall' vs 'validator stall' (the design's own not_measurable #4) remains undecided — which is exactly the state today plus an error string. The SIGSTOP negative control does not validate what it claims: SIGSTOP freezes every thread (unlike the one-lock/one-worker mechanisms listed), returns .missing pprof by construction, cannot exercise the deaf-listener / fd-cap / SQ-full classes at all, and the 25 s variant trips the existing I-HANG oracle (4 x 5 s probes 2 s apart, liveness.go:243) into a hard fail that competes with the predicted I-H2C-HANG dossier for the same 30 s forensics budget. So passing the control proves the ring gets written, not that the next event will be root-causable. Minor premise errors: the weekend tier is now 64 cells/4 engines (adaptive) since 2026-09-13, so '3 engines / EXPECT_CELLS=6' is stale; the LocalAddr join key joins against nothing because the refapp logs nothing per-conn (io.Discard recovery logger, no engine per-conn log).
Change to the design: 1) Drop the 18 h directed run as a discriminator. State the rate honestly (one event => 95% interval [0.025,5.6]x) and pre-register a stopping rule on EVENTS, not hours: keep the instrumentation on in every routine soak and stop when >=10 attributed events have accumulated per counter, or when a pre-registered denominator (e.g. 5M h2c reads and 2M WS handshakes on the same config) has 0 events — only that bounds the rate below the v1.5.11 point estimate with power. A directed cell adds nothing a routine soak does not give for free once capture is unconditional. 2) Make the mechanism observable independent of the SUT: add a loopback net/http side-listener in debugvars (ephemeral port announced in the ready banner, the #493 diag pattern) serving /debug/vars + /debug/pprof, and point BOTH the 1 Hz property poll and the dossier at it; keep the engine-served path too, so 'engine path timed out, side path answered' is itself the process-wide-stall signal. 3) Take the dossier INSIDE the stall, not after: have the h2c/WS walkers raise a pre-fail probe when a read has gone >1 s (WS) / >3 s (h2c) with no bytes — goroutine dump via the side listener, ss -tnp/ss -ltnp (listen Recv-Q + the walker's LocalAddr on the server side: recvq>0 = accepted-unread, absent = never accepted — the discriminator that actually split 16/26 vs 10/26 in #493), and a /debug/vars conn-table snapshot to join the walker's LocalAddr against; fire per event (drop alertedCounters for these two, bounded to e.g. 16 dossiers per cell). 4) Replace SIGSTOP with mechanism-shaped fault injections behind an env gate in the refapp: (a) a route that holds one store shard's lock for N s, (b) an inline worker sleep of N s, (c) an accept-side drop (ulimit -n to force EMFILE, or an env that skips the accept re-arm once) — and assert that each yields the DIFFERENT expected classification (a/b: side-listener dump shows workers in Mutex.Lock / off the syscall, ss shows backlog; c: engine_error_count steps, listen Recv-Q>0 with idle workers). Keep SIGSTOP only as the host-wide case and set its duration <20 s so it does not collide with I-HANG. 5) Fix the std control: WS (middleware-free) is the engine-loop control; the h2c GET / is not, and must be compared per path. 6) Update the premises: 4 engines / 64 cells; include adaptive or exclude it explicitly; derive EXPECT_CELLS from the selection.
Review 2: refuted
Why: The design's rate/mechanism leg (B) and its main discriminators measure the wrong thing; the attributability leg (A) is fine in intent but the instruments it names cannot deliver the signatures it says will separate TRUE from FALSE.
The h2c channel is calibrated on an instrument that no longer exists. The v1.5.11 event was a timeout at the then 10 s read budget; the budget is now 20 s (validation/h2c.go:185-198, feat: per-handler async dispatch (sync/async per route + group, all engines × all protocols) #302). The strongest prior the design itself cites, celeris#493, was stalls of 1.5-8 s once a minute (issue text: "stalls of 1.5-8 s happen once a minute at the refapp's start second", goroutine-dump proven). A io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493-residual stall can NEVER produce an h2c_hang at a 20 s budget: the read simply returns late as declined. So the design's TRUE prediction "~6 h2c hangs per 6 h cell at ~1/120k, read elapsed ~20000 ms" extrapolates a 10 s-budget rate to a 20 s-budget walker, and its FALSE reading "0 h2c events bounds the rate below ~1/1M / fix(engine): clear the hand-off queue slots so a drained conn can be collected #574-era changes removed it" would be exactly what the mechanism being present looks like through the new instrument. Only the WS channel (still 2 s, ws.go fireWSTorture) is comparable across the two soaks. The walker records nothing about slow-but-answered reads (h2c.go:286-322 only branches on err/n==0), so the 1-19 s tail the prior predicts is invisible by construction.
Whole-server stalls >= 20 s are already a different oracle. I-HANG is 4 consecutive 5 s /healthz timeouts at 2 s spacing (liveness.go:243-247) = ~20 s, the same instant the h2c read budget expires, and I-HANG is a HARD fail: writeIncidentDossier + captureForensics then cancel() and abort (runner.go:640-660; only RecordOnly incidents keep the cell running, 629-638). Consequences: (a) the proposed 25 s SIGSTOP negative control trips I-HANG at ~T+20-22 s and aborts the cell — it validates the hang oracle, not the h2c capture path; (b) an h2c_hang_timeout in a cell that did NOT also I-HANG (the v1.5.11 case) implies a PARTIAL stall (some workers) or a deaf listener, i.e. exactly the class a whole-process SIGSTOP does not model. The control injects the wrong fault class for the surviving hypothesis space.
The cross-slice time-join is not measurable with the proposed changes. The deps: bump the all-go-deps group across 3 directories with 5 updates #327 SSE ring the design models itself on is earlyErrs []string with no timestamp (sse.go:73-79, 138-144), and there is no adversarial ring at all (adversarial.go: bare atomic counters). The code_changes_needed add timestamped rings only for h2c/ws, so "join ts against the SSE/adversarial rings" — the headline TRUE-vs-FALSE discriminator — has no data on the other side of the join.
The validator-side confound is measured with the wrong metric. A 2 s WS timeout cannot come from a Go GC pause; the plausible client-side stall is scheduler/CPU starvation of the validator process sharing the host, or a host-wide freeze. runtime/metrics GC pause would read ~0 in exactly the case it is meant to flag, and no control ever injects a validator-side stall, so the record's client/server attribution is never validated in that direction.
Minor: the SIGSTOP(3 s) assertion "ts within [T, T+3 s]" is wrong even for the intended case — WS failures are recorded at deadline expiry, i.e. [T+0, T+5 s]; the weekend workflow now expects 64 cells (4 engines incl. adaptive), not 48.
Change to the design: Keep the per-event rings, stderr tail, engine_error_count export and ready_at (leg A), but change what the run measures and how the controls are built:
Measure the read-latency DISTRIBUTION, not only the timeout tail. In h2c.go and ws.go record every fire's dial/write/read elapsed into fixed buckets (<100 ms, <1 s, <2 s, <5 s, <10 s, <20 s, timeout) plus a bounded ring (16-64) of every fire whose read exceeded 1 s (ts, elapsed, outcome, err, local_addr). This makes the v1.5.11 10 s-budget event comparable (it is the >=10 s bucket), makes a io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493-residual 1.5-8 s stall visible as a once-per-minute burst in the 1-10 s buckets regardless of walker budget, and turns "0 h2c_hang" into a statement about the whole tail rather than about one threshold. Compute the mod-60 phase histogram over these slow-read events, not over hangs.
Replace the fire()-triggered dossier with a stall-triggered sampler that runs continuously in the cell: sample ss -ltn (Recv-Q on the refapp's LISTEN sockets) every 200 ms and, on backlog >= 2 for >= 400 ms, take goroutine.pprof + /proc//stack + the ss snapshot, rate-limited to N per cell and independent of tallyCB dedupe. This is the instrument that root-caused io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493 and it captures the state INSIDE a 1.5-8 s stall, which no post-event dossier can. Per-event dossiers via fire() may stay as a fallback but must not be the discriminator.
Fix the negative controls to the surviving hypothesis classes, not whole-process freeze: (a) partial stall — a validation-only refapp flag (e.g. -fault=store-lock:<dur>@<period>) that holds one session-store shard lock for 3 s (and once for 8 s) on a 60 s period, reproducing io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493's shape: healthz stays alive (no I-HANG), WS fails at 2 s, h2c reads land in the 2-10 s buckets, backlog sampler fires, phase histogram peaks; (b) deaf-listener/fd-cap — run the refapp under ulimit -n small enough that accepts hit EMFILE, asserting engine_error_count steps and the epoll fd-cap Warn line lands in refapp_stderr_tail; (c) keep SIGSTOP only at 3 s (below I-HANG), asserting WS timeouts in [T, T+5 s], and drop the 25 s SIGSTOP (it measures I-HANG). (d) validator-side control: SIGSTOP the VALIDATOR for 3 s and assert the events it produces are flagged client-side.
Replace the validator GC-pause field with a heartbeat-skew monitor in the validator (a 100 ms ticker whose observed gap is recorded; any gap > 500 ms in the +/-5 s window of an event is stamped into the record) and/or runtime/metrics /sched/latencies:seconds; this is what actually discriminates a client-process or host-wide stall from a server stall.
Give the SSE and adversarial slices the same timestamped event ring (ts, err, elapsed, local_addr) so the cross-slice time-join exists on both sides; alternatively one shared per-cell event log (JSONL, bounded) that all four walkers append to.
Restate the rate expectation for the directed run per channel: WS (unchanged 2 s budget) inherits the v1.5.11/v1.5.9 rate and is the only counter whose 0-in-N is a bound on the old rate; h2c is judged on the >=1 s slow-read ring and the >=10 s bucket, never on h2c_hang alone. Set VALIDATE_GATE_EXPECT_CELLS to the actual cell count of the scoped dispatch (the workflow now assumes 64, four engines).
Filed from the release-readiness sweep of 2026-09-13 under the standing rule that a closure counts only when it is MEASURED. Design produced by a read-only design pass and corrected by two adversarial reviews; the corrections are binding and listed last.
Closure under test
v1.5.11 24h soak (probatorium run 34003151551, celeris v1.5.11 pins, probatorium f8e5d3c): auth_session_ratelimit/iouring/arm64 tier_1.h2c_churn.h2c_hang=1 (cause=timeout, max_elapsed=10000 ms) and auth_session_ratelimit/epoll/arm64 tier_1.ws_torture.ws_handshake_fail=1 (pre-cause-split); neither root-caused; 0 recurrences in the 96 cells of soaks 34368601903 and 34616620237.
What the closure claims
NOT closable as root-caused, and NOT closable as "gone". The two original events are permanently unattributable: the artifact kept only the counters (verified: the msr1 cell-03 directory in the downloaded v1.5.11 artifact held nothing but incidents/20260906-043231-I-MEM-1; no tally sidecar, no error string, no timestamp, no engine log). 0-in-96 later cells is statistically compatible with the observed per-cell rate (~1 h2c hang per ~120k reads, ~1 WS failure per ~47.6k handshakes at concurrency 50 — the memory's own "0-in-55k and 1-in-55k are the same rate" lesson applies). What CAN be closed by measurement is the apparatus claim "the next occurrence will be root-causable from the artifact alone" — that claim is FALSE today, and the measurement below turns it TRUE and verifies it with a fault-injected negative control before the next soak. Separately, a directed 6 h cell can force the rate into a measurable regime (~6 events per cell per counter if the observed rate holds).
What was argued rather than measured
GET /(h2c.go:337-358) that hits session+ratelimit (main.go:215-245;/is not routed, so 404 or 429 = "declined"). Correct that it is not an upgrade path, but stall-vs-deaf-listener-vs-server-header-timeout was never discriminated: h2c.go:185-194 says in so many words that at the then-10 s read budget "it is not recoverable from the artifact whether the server stalled or merely hit its own header deadline" (ReadHeaderTimeout default 10 s, config.go:92-102; enforced in iouring worker.go:1479/2796 and epoll loop.go:297). The budget was raised to 20 s in feat: per-handler async dispatch (sync/async per route + group, all engines × all protocols) #302 (merged 2026-09-07, AFTER this soak).Measurement
Observable
Two things, in order. (A) Attributability: for every h2c_hang and ws_handshake_fail event the cell document must carry a per-event record (bounded ring of 16, modelled on #327's sse_early_errs at sse.go:73-79,107-145 / schema.go:618-623): UTC ts (RFC3339Nano), ms since refapp ready, walker seed+id, churn/torture mode, cause class, verbatim err.Error(), non-101 status line (WS), walker conn.LocalAddr() (the engine<->client join key from celeris#562), THIS event's dial/write/read durations and bytes read, plus validator-side GC pause since last fire (runtime/metrics) — and, alongside it, an incident dossier taken within 2 s (goroutine.pprof, /proc//stack, fd list, dmesg), the refapp stderr ring flushed to /refapp_stderr_tail.txt, celeris.engine_error_count added to /debug/vars and the 1 Hz series, and ready_at in the cell doc. (B) Rate/mechanism: count of events per 6 h directed cell per counter, their timestamps, and the series columns accepted/active/engine_error_count in the +/-10 s window around each event.
Where it runs
(A1) Unit, probatorium validation/ under -race on Linux: extend newFakeH2CServer (h2c_test.go:89-113) and newFakeWSServer (ws_test.go:163-217) with policies {silent 25 s, accept-then-close, RST, slow-101 at 3 s} and assert the per-event record fields. (A2) Apparatus negative control in a Linux container with --security-opt seccomp=unconfined --ulimit memlock=-1 (the imem1 rig at scratchpad/imem1/{drive.sh,run.sh}): one auth_session_ratelimit io_uring cell, concurrency 50, 20 min; at a known instant T a sidecar sends SIGSTOP to the refapp for 3 s (and once for 25 s), then SIGCONT. (B) Probatorium cluster weekend tier, scoped: VALIDATE_MATRIX=1 VALIDATE_MATRIX_REFAPPS=auth_session_ratelimit VALIDATE_MATRIX_ENGINES=iouring,epoll,std SOAK_DURATION=18h VALIDATE_TARGET=both VALIDATE_CONCURRENCY=50 -> cellBudget = 18h/3 = 6 h per cell (cmd/validator/matrix.go:403), both hosts in parallel (VALIDATE_PARALLEL=1). The dispatch currently exposes only duration/resume_from (matrix-weekend-tier.yml:15-30), so this needs either a workflow input or a
mage Soakfrom the runner host; VALIDATE_GATE_EXPECT_CELLS must be 6 (not 48) for that run.Procedure
go test -race ./validation/... ./report/...on Linux. 2. Run the fault-injected container cell: verify (a) the 3 s SIGSTOP produces ws_handshake_fail_timeout events whose ts is within [T, T+3 s], h2c_hang=0 (3 s < 20 s), sse early closes/timeouts at the same T, seriesacceptedflat for 3 rows, a dossier dir incidents/-I-WS-HANDSHAKE with goroutine.pprof, refapp_stderr_tail.txt present; (b) the 25 s SIGSTOP additionally produces h2c_hang_timeout with per-event read elapsed ~20000 ms and I-H2C-HANG dossier; (c) with no injection over 20 min all new rings are empty and the gate still PASSes (guard test like deps: bump the all-go-deps group across 3 directories with 5 updates #327's fails if the rings are dropped from schema or gate). 3. Dispatch the scoped 18 h soak (6 h x 3 engines per host). 4. From the artifact: tally events per cell per counter; for each event join ts against the series and against the SSE/adversarial rings; compute (ts - ready_at) mod 60 s; read the dossier's goroutine.pprof for the worker goroutines' state; read refapp_stderr_tail for engine Warn/Error lines. 5. Compare iouring vs epoll vs std per host.Expected if the claim is TRUE
TRUE = a server-side mechanism exists at the observed rate. In each 6 h iouring/epoll cell expect ~6 h2c hangs (5 walkers x 10/s x 21600 s x 2/3 non-RST reads = ~720k reads at ~1/120k) and ~6 WS failures (2 walkers x 6.67/s x 21600 s = ~288k at ~1/47.6k) with 95% bands roughly [2,12]; events CLUSTER in time across slices (h2c, WS, SSE early-close, adv_hang_until_timeout within the same few seconds); h2c per-event read elapsed ~20000 ms with cause=timeout and WS ~2000 ms timeout for a stall, OR small-elapsed eof/reset for a drop; series
acceptedflat (oractivespiking) across the event window; goroutine.pprof shows worker goroutines off the io_uring_enter/epoll_wait syscall (blocked on a lock or running a sweep) for the stall class, or the listen worker idle with no accept SQE for the deaf-listener class; celeris.engine_error_count steps at the event for the fd-cap/EMFILE class; if #493-residual, (ts - ready_at) mod 60 clusters within a few seconds of one phase; std shows 0 (no worker event loop). The error string discriminates further: 'i/o timeout' = never answered; 'EOF' at ~10000 ms = server ReadHeaderTimeout closed an accepted-but-unread conn; 'connection reset by peer' small elapsed = fd dropped at accept.Expected if the claim is FALSE
FALSE = apparatus/transport artefact (the ~1/9000 dial-loss class the SSE analysis already found on all engines). Events are uniformly spread, do not coincide with SSE/adv events,
acceptednever flattens, goroutine.pprof is unremarkable, engine_error_count does not move, the phase histogram is uniform, the new h2c/ws dial_fail and write_fail counters show a comparable rate on std, and the error strings are client-side (dial i/o timeout, write: broken pipe) with validator GC pause > 0 at the instant. Also possible: 0 events in 6 h x 3 cells per host (~2.2M h2c reads, ~860k WS handshakes total) which bounds the rate below ~1/1M and says the v1.5.11 pair was a rate that #574/#572-era changes removed — that is a measured statement, unlike today's 0-in-96.Negative controls
Three: (1) the fault-injected SIGSTOP cell above (known instant, known duration) must produce exactly the predicted counters/records/dossier with timestamps within 2 s of T — this proves the capture path works before it is relied upon; without injection the same cell must produce empty rings and a PASS (proves no false positives; run it also with the rings assertion inverted to confirm the guard test fails when the fields are removed, per RULE: always run the negative control); (2) the std-engine 6 h cell in the same dispatch: std has no worker event loop and no fd-cap table, so a server-side stall/deaf-listener mechanism must NOT appear there while an apparatus artefact appears at the same rate; (3) unit fake-server policies assert the record content for each cause class so a future refactor cannot silently drop the error string again (the #327 guard-test pattern).
Cost
Unit tests: ~1-2 min per package under -race. Container fault-injection: ~30 min wall (2 x 20 min cells) on any Linux box with seccomp=unconfined. Directed cluster run: 18 h wall on both hosts in parallel (6 cells total, 3 per host), exclusive cluster time; if 0 events, one repeat at VALIDATE_CONCURRENCY=100 (h2c 10 walkers, WS 5) for another 18 h before bounding the rate — note that changes the load profile so its I-MEM verdicts are not comparable to the standard soak. msr1 is the runner host: nothing may touch it during the run. Code: ~1 day incl. schema bump and tests.
Not measurable, and why
(1) The two original v1.5.11 events: no error string, timestamp, stderr tail or dossier exists anywhere (artifact verified; /tmp/soak3 is empty; the run artifact 340 KB is downloadable but contains only the run document and I-MEM-1 dossiers per the memory's own audit). They stay unattributed forever. (2) Stalls shorter than the 2 s tallyCB interval (runner.go:1001) will have a dossier taken after the stall ended; only the per-event record + 1 Hz series catch them, and a <1 s stall cannot fail either walker anyway (WS budget 2 s). (3) Kernel accept-queue occupancy at the instant (ss -ltn / /proc/net/tcp) is not captured by forensics.go; 'deaf listener' vs 'worker stalled after accept' is inferred from elapsed+EOF patterns and goroutine state, not observed directly — adding an
ss -ltnpsnapshot to the dossier would close that but is a further change. (4) The validator shares the host and loopback with the refapp, so a validator-process scheduler stall is indistinguishable from a server stall from the walker's clock alone; the record's validator GC-pause field and the cross-slice clustering bound this but cannot exclude a host-wide stall (both processes freeze together). (5) A rate below ~1/1M per counter is not resolvable in one 6 h cell; only repeated directed runs can push the bound down.Code changes needed
probatorium (all under validation/ unless noted): (1) h2c.go: recordHang(err, elapsed) -> recordHang(ev h2cEvent) with a mutex-guarded ring of 16 {ts, since_ready_ms, walker, mode, cause, err, local_addr, dial_ms, write_ms, read_ms, n_read, validator_gc_pause_ns}; count dial/write failures as h2c_dial_fail / h2c_write_fail (informational, not gated); snapshot() emits h2c_hang_events. (2) ws.go: same for recordHandshakeFail (+ status line text), ws_dial_fail / ws_write_fail; keep the 2 s wsMaxHold but split the dial deadline from the handshake read deadline so the record can say which leg expired. (3) runner.go Tier1Summary projection + report/schema.go: new H2CHangEvents / WSHandshakeFailEvents []struct fields next to SSEEarlyErrs (schema.go:618-623), schema minor bump; report/gate.go causeSuffix prints the first event's err/elapsed (informational). (4) runner.go tallyCB: fire(properties.IH2CHang.ID, ..., snap.H2CChurn.Hang > 0) and fire(properties.IWSHandshake.ID, ..., snap.WSTorture.HandshakeFail > 0) with RecordOnly=true so the cell runs on but a dossier is written; add the two Specs to properties/i_tier1_walker.go and registry.go. (5) liveness.go: expose the stderr ring via a tailSnapshot() on livenessTally; runner.go writeIncidentDossier writes refapp_stderr_tail.txt into the dossier and tier1 end-of-cell writes /refapp_stderr_tail.txt unconditionally (80 lines, bounded); optionally tee the refapp's merged output to /refapp.log with a size cap in remote/local.go. (6) refapp/internal/debugvars: export celeris.engine_error_count from EngineMetrics.ErrorCount; checker/poll.go Snapshot field; series.go add the column; record the ready-banner time as ready_at in the cell doc (runner.go/matrix.go). (7) .github/workflows/matrix-weekend-tier.yml: dispatch inputs refapps/engines mapped to VALIDATE_MATRIX_REFAPPS/ENGINES and VALIDATE_GATE_EXPECT_CELLS derived from the selection. (8) Tests: h2c_class_test.go / ws_test.go per-event assertions for silent/close/RST/slow policies; a gate guard test that fails if the event rings vanish from schema or the gated totals stop being gated. celeris (optional but raises signal): bump errCount when prepareAccept hits a full SQ ring (engine/iouring/worker.go:3586-3591 currently only sets acceptRearmPending) and expose fdCapDrops (engine/epoll/loop.go:129-135) through EngineMetrics.ErrorCount, so the exported counter moves on exactly the suspected deaf-listener paths.
Corrections from the adversarial review (binding)
Review 1: refuted
Why: Part (B) cannot discriminate its two hypotheses at the sample size proposed. The whole basis is ONE h2c event and ONE WS event; a single Poisson observation has a 95% rate interval of roughly [0.025, 5.6] x the point estimate, so the 'expected_if_true' band '[2,12] events per 6 h cell' is fiction — the honest expectation for a 6 h cell (720k h2c reads / 288k WS handshakes) is anywhere from 0.05 to ~11 events, and the pooled data already in hand (1 in 3 same-config cells, i.e. ~1/360k h2c and ~1/143k WS; or 1 in 5.8M h2c / 1 in ~1M WS if uniform across cells) gives lambda ~2 with P(0)=0.14, or lambda 0.04-0.3 with P(0)>=0.74. A 0-event result is therefore compatible with 'mechanism still present', 'mechanism removed' and 'apparatus artefact' alike, and the design explicitly plans to read 0 events as the FALSE branch ('bounds the rate below ~1/1M ... says the v1.5.11 pair was removed') when it bounds nothing the 96 routine cells have not already bounded. The std 'negative control' is powerless for the same reason: at lambda 0.04-0.3 std shows 0 under BOTH hypotheses (P>=0.74), so 'must NOT appear on std' is satisfied by chance. It is also mis-specified: the h2c GET / goes through session+ratelimit on std too (main.go:214-245 skips only /ws,/events), so a lock-hold stall WOULD show on std for h2c and not for WS. Part (A)'s key observable does not survive contact with the capture path: /debug/vars and /debug/pprof are served THROUGH the engine under test (debugvars.go:217 srv.Pre(v.Handler(), pprof.New())), so for the process-wide worker-freeze class the design says expected_if_true hinges on ('goroutine.pprof shows workers off the io_uring_enter/epoll_wait syscall'), the 500 ms poll and the 5 s pprof fetch both fail during the stall and the dossier holds goroutine.pprof.missing; celeris#493 was only proven because the diag build had a pprof SIDE-listener and a 200 ms ss sampler — neither is in this design (ss is explicitly deferred as 'a further change'). Even with a side listener, alertedCounters fires the dossier once per counter per cell and only after the walker's budget expired (>=2 s WS, >=20 s h2c, +<=2 s tallyCB), so for the observed 1.5-8 s stall class the h2c dossier is always post-stall and the WS one is post-stall for anything under ~4 s: the record says 'timeout', the dossier shows a healthy process, and 'server stall' vs 'host-wide stall' vs 'validator stall' (the design's own not_measurable #4) remains undecided — which is exactly the state today plus an error string. The SIGSTOP negative control does not validate what it claims: SIGSTOP freezes every thread (unlike the one-lock/one-worker mechanisms listed), returns .missing pprof by construction, cannot exercise the deaf-listener / fd-cap / SQ-full classes at all, and the 25 s variant trips the existing I-HANG oracle (4 x 5 s probes 2 s apart, liveness.go:243) into a hard fail that competes with the predicted I-H2C-HANG dossier for the same 30 s forensics budget. So passing the control proves the ring gets written, not that the next event will be root-causable. Minor premise errors: the weekend tier is now 64 cells/4 engines (adaptive) since 2026-09-13, so '3 engines / EXPECT_CELLS=6' is stale; the LocalAddr join key joins against nothing because the refapp logs nothing per-conn (io.Discard recovery logger, no engine per-conn log).
Change to the design: 1) Drop the 18 h directed run as a discriminator. State the rate honestly (one event => 95% interval [0.025,5.6]x) and pre-register a stopping rule on EVENTS, not hours: keep the instrumentation on in every routine soak and stop when >=10 attributed events have accumulated per counter, or when a pre-registered denominator (e.g. 5M h2c reads and 2M WS handshakes on the same config) has 0 events — only that bounds the rate below the v1.5.11 point estimate with power. A directed cell adds nothing a routine soak does not give for free once capture is unconditional. 2) Make the mechanism observable independent of the SUT: add a loopback net/http side-listener in debugvars (ephemeral port announced in the ready banner, the #493 diag pattern) serving /debug/vars + /debug/pprof, and point BOTH the 1 Hz property poll and the dossier at it; keep the engine-served path too, so 'engine path timed out, side path answered' is itself the process-wide-stall signal. 3) Take the dossier INSIDE the stall, not after: have the h2c/WS walkers raise a pre-fail probe when a read has gone >1 s (WS) / >3 s (h2c) with no bytes — goroutine dump via the side listener,
ss -tnp/ss -ltnp(listen Recv-Q + the walker's LocalAddr on the server side: recvq>0 = accepted-unread, absent = never accepted — the discriminator that actually split 16/26 vs 10/26 in #493), and a /debug/vars conn-table snapshot to join the walker's LocalAddr against; fire per event (drop alertedCounters for these two, bounded to e.g. 16 dossiers per cell). 4) Replace SIGSTOP with mechanism-shaped fault injections behind an env gate in the refapp: (a) a route that holds one store shard's lock for N s, (b) an inline worker sleep of N s, (c) an accept-side drop (ulimit -n to force EMFILE, or an env that skips the accept re-arm once) — and assert that each yields the DIFFERENT expected classification (a/b: side-listener dump shows workers in Mutex.Lock / off the syscall, ss shows backlog; c: engine_error_count steps, listen Recv-Q>0 with idle workers). Keep SIGSTOP only as the host-wide case and set its duration <20 s so it does not collide with I-HANG. 5) Fix the std control: WS (middleware-free) is the engine-loop control; the h2c GET / is not, and must be compared per path. 6) Update the premises: 4 engines / 64 cells; include adaptive or exclude it explicitly; derive EXPECT_CELLS from the selection.Review 2: refuted
Why: The design's rate/mechanism leg (B) and its main discriminators measure the wrong thing; the attributability leg (A) is fine in intent but the instruments it names cannot deliver the signatures it says will separate TRUE from FALSE.
The h2c channel is calibrated on an instrument that no longer exists. The v1.5.11 event was a timeout at the then 10 s read budget; the budget is now 20 s (validation/h2c.go:185-198, feat: per-handler async dispatch (sync/async per route + group, all engines × all protocols) #302). The strongest prior the design itself cites, celeris#493, was stalls of 1.5-8 s once a minute (issue text: "stalls of 1.5-8 s happen once a minute at the refapp's start second", goroutine-dump proven). A io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493-residual stall can NEVER produce an h2c_hang at a 20 s budget: the read simply returns late as
declined. So the design's TRUE prediction "~6 h2c hangs per 6 h cell at ~1/120k, read elapsed ~20000 ms" extrapolates a 10 s-budget rate to a 20 s-budget walker, and its FALSE reading "0 h2c events bounds the rate below ~1/1M / fix(engine): clear the hand-off queue slots so a drained conn can be collected #574-era changes removed it" would be exactly what the mechanism being present looks like through the new instrument. Only the WS channel (still 2 s, ws.go fireWSTorture) is comparable across the two soaks. The walker records nothing about slow-but-answered reads (h2c.go:286-322 only branches on err/n==0), so the 1-19 s tail the prior predicts is invisible by construction.Whole-server stalls >= 20 s are already a different oracle. I-HANG is 4 consecutive 5 s /healthz timeouts at 2 s spacing (liveness.go:243-247) = ~20 s, the same instant the h2c read budget expires, and I-HANG is a HARD fail: writeIncidentDossier + captureForensics then cancel() and abort (runner.go:640-660; only RecordOnly incidents keep the cell running, 629-638). Consequences: (a) the proposed 25 s SIGSTOP negative control trips I-HANG at ~T+20-22 s and aborts the cell — it validates the hang oracle, not the h2c capture path; (b) an h2c_hang_timeout in a cell that did NOT also I-HANG (the v1.5.11 case) implies a PARTIAL stall (some workers) or a deaf listener, i.e. exactly the class a whole-process SIGSTOP does not model. The control injects the wrong fault class for the surviving hypothesis space.
The dossier cannot discriminate the prior-supported class. The design routes dossiers through tallyCB fire(), which is deduped "AT MOST ONCE per counter per run" (runner.go:922-940) — one dossier per cell, not one per event, so "events CLUSTER" and "read each dossier's goroutine state" are unmeasurable for events 2..n. Worse, for the 1.5-8 s io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493 class the WS event is recorded 2 s after handshake start and the tick adds up to 2 s more, so goroutine.pprof lands after the stall ended and is "unremarkable" — which the design lists as the FALSE signature. What actually proved io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493 was a 200 ms
ss -ltnbacklog sampler triggering goroutine dumps INSIDE the stall (memory absolute-gate-v159-release.md:160-171, stall-probe3); the design demotes exactly that instrument to "not measurable / further change".The cross-slice time-join is not measurable with the proposed changes. The deps: bump the all-go-deps group across 3 directories with 5 updates #327 SSE ring the design models itself on is
earlyErrs []stringwith no timestamp (sse.go:73-79, 138-144), and there is no adversarial ring at all (adversarial.go: bare atomic counters). The code_changes_needed add timestamped rings only for h2c/ws, so "join ts against the SSE/adversarial rings" — the headline TRUE-vs-FALSE discriminator — has no data on the other side of the join.The validator-side confound is measured with the wrong metric. A 2 s WS timeout cannot come from a Go GC pause; the plausible client-side stall is scheduler/CPU starvation of the validator process sharing the host, or a host-wide freeze. runtime/metrics GC pause would read ~0 in exactly the case it is meant to flag, and no control ever injects a validator-side stall, so the record's client/server attribution is never validated in that direction.
Minor: the SIGSTOP(3 s) assertion "ts within [T, T+3 s]" is wrong even for the intended case — WS failures are recorded at deadline expiry, i.e. [T+0, T+5 s]; the weekend workflow now expects 64 cells (4 engines incl. adaptive), not 48.
Change to the design: Keep the per-event rings, stderr tail, engine_error_count export and ready_at (leg A), but change what the run measures and how the controls are built:
Measure the read-latency DISTRIBUTION, not only the timeout tail. In h2c.go and ws.go record every fire's dial/write/read elapsed into fixed buckets (<100 ms, <1 s, <2 s, <5 s, <10 s, <20 s, timeout) plus a bounded ring (16-64) of every fire whose read exceeded 1 s (ts, elapsed, outcome, err, local_addr). This makes the v1.5.11 10 s-budget event comparable (it is the >=10 s bucket), makes a io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493-residual 1.5-8 s stall visible as a once-per-minute burst in the 1-10 s buckets regardless of walker budget, and turns "0 h2c_hang" into a statement about the whole tail rather than about one threshold. Compute the mod-60 phase histogram over these slow-read events, not over hangs.
Replace the fire()-triggered dossier with a stall-triggered sampler that runs continuously in the cell: sample
ss -ltn(Recv-Q on the refapp's LISTEN sockets) every 200 ms and, on backlog >= 2 for >= 400 ms, take goroutine.pprof + /proc//stack + the ss snapshot, rate-limited to N per cell and independent of tallyCB dedupe. This is the instrument that root-caused io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493 and it captures the state INSIDE a 1.5-8 s stall, which no post-event dossier can. Per-event dossiers via fire() may stay as a fallback but must not be the discriminator.Fix the negative controls to the surviving hypothesis classes, not whole-process freeze: (a) partial stall — a validation-only refapp flag (e.g.
-fault=store-lock:<dur>@<period>) that holds one session-store shard lock for 3 s (and once for 8 s) on a 60 s period, reproducing io_uring worker event loops freeze once a minute: session store sweep holds shard locks for seconds while the inline session Set blocks every worker (24h soak gate failure) #493's shape: healthz stays alive (no I-HANG), WS fails at 2 s, h2c reads land in the 2-10 s buckets, backlog sampler fires, phase histogram peaks; (b) deaf-listener/fd-cap — run the refapp underulimit -nsmall enough that accepts hit EMFILE, asserting engine_error_count steps and the epoll fd-cap Warn line lands in refapp_stderr_tail; (c) keep SIGSTOP only at 3 s (below I-HANG), asserting WS timeouts in [T, T+5 s], and drop the 25 s SIGSTOP (it measures I-HANG). (d) validator-side control: SIGSTOP the VALIDATOR for 3 s and assert the events it produces are flagged client-side.Replace the validator GC-pause field with a heartbeat-skew monitor in the validator (a 100 ms ticker whose observed gap is recorded; any gap > 500 ms in the +/-5 s window of an event is stamped into the record) and/or runtime/metrics
/sched/latencies:seconds; this is what actually discriminates a client-process or host-wide stall from a server stall.Give the SSE and adversarial slices the same timestamped event ring (ts, err, elapsed, local_addr) so the cross-slice time-join exists on both sides; alternatively one shared per-cell event log (JSONL, bounded) that all four walkers append to.
Restate the rate expectation for the directed run per channel: WS (unchanged 2 s budget) inherits the v1.5.11/v1.5.9 rate and is the only counter whose 0-in-N is a bound on the old rate; h2c is judged on the >=1 s slow-read ring and the >=10 s bucket, never on h2c_hang alone. Set VALIDATE_GATE_EXPECT_CELLS to the actual cell count of the scoped dispatch (the workflow now assumes 64, four engines).