Summary
Across the last six nightly validation runs (288 cells, ~1.04M h2c requests), the h2c_hang Tier 1 invariant counter fires only on the io_uring engine — never on epoll or std, under an identical workload.
| engine |
cells |
cells with h2c_hang |
rate |
h2c_sent total |
| epoll |
96 |
0 |
0.0% |
345,312 |
| iouring |
96 |
14 |
14.6% |
345,213 |
| std |
96 |
0 |
0.0% |
345,312 |
h2c_hang is a documented "must be zero" Tier 1 invariant (report/diff.go:invariantCounters).
Why this is a real wedge, not measurement noise
The counter has been deliberately de-noised in probatorium (validation/h2c.go):
h2c_intentional_rst was split out of hang so the workload's own RST-before-read no longer pollutes it — the comment notes a 3-day soak once showed h2c_hang=317K that "was almost entirely this workload noise".
- The read budget was raised 2s → 10s specifically to eliminate ~18K false positives per soak from slow-but-correct refapps.
So a h2c_hang today means: the server accepted the connection, received a valid h2c upgrade preamble, and produced zero bytes for a full 10 seconds.
Distribution — engine-specific, not arch- or app-specific
- by engine:
{iouring: 14} — 14 of 14
- by arch:
{amd64: 7, arm64: 7} — evenly split, so not an arm64 issue
- by refapp: spread over 7 different refapps (
driver_redis 4, kitchen_sink 3, auth_session_ratelimit 2, driver_memcached 2, observability, static_swagger_proxy, auth_jwt_csrf)
Every affected cell shows the same shape: h2c_sent=3597, h2c_upgraded=0, h2c_declined~2463, h2c_hang=1..4.
Overall rate ≈ 1 hang per ~24,700 h2c requests, present in every run:
run cells non-zero invariant counters
33546772656 48 MED:h2c_hang=1 LOW:sse_handshake_fail=1
33482366357 48 MED:h2c_hang=2
33243914887 48 MED:h2c_hang=2 LOW:sse_handshake_fail=1
33071374027 48 MED:h2c_hang=3 LOW:sse_handshake_fail=1
32924910721 48 MED:h2c_hang=4
32803245801 48 MED:h2c_hang=2
Related: sse_handshake_fail is also io_uring-only
3 events across the same six runs, all io_uring (observability, kitchen_sink, auth_jwt_csrf — all arm64). Possibly the same underlying accept/upgrade path; noting it here rather than opening a second issue until the cause is known.
Good news
No HIGH-severity invariant has ever fired across all 288 cells: adv_wrong_accepted, h2c_crashed and ws_accepted_bad_frame are all clean. Nothing is crashing or accepting malformed input — this is a liveness/wedge issue, not a correctness one.
Why it went unnoticed
Reported separately in goceleris/probatorium — the short version is that h2c_hang is MED severity, MED only fails the gate under VALIDATE_DIFF_STRICT (never set in any workflow), and the cross-arch diff only flags a counter when it is zero on exactly one side. Because io_uring hangs on both arches, most nights it is either symmetric (invisible) or coincidentally asymmetric (reported, but non-fatal).
Evidence gathered from nightly artifacts matrix-nightly-results-{33546772656,33482366357,33243914887,33071374027,32924910721,32803245801}.
Summary
Across the last six nightly validation runs (288 cells, ~1.04M h2c requests), the
h2c_hangTier 1 invariant counter fires only on the io_uring engine — never on epoll or std, under an identical workload.h2c_hangh2c_senttotalh2c_hangis a documented "must be zero" Tier 1 invariant (report/diff.go:invariantCounters).Why this is a real wedge, not measurement noise
The counter has been deliberately de-noised in probatorium (
validation/h2c.go):h2c_intentional_rstwas split out ofhangso the workload's own RST-before-read no longer pollutes it — the comment notes a 3-day soak once showedh2c_hang=317Kthat "was almost entirely this workload noise".So a
h2c_hangtoday means: the server accepted the connection, received a valid h2c upgrade preamble, and produced zero bytes for a full 10 seconds.Distribution — engine-specific, not arch- or app-specific
{iouring: 14}— 14 of 14{amd64: 7, arm64: 7}— evenly split, so not an arm64 issuedriver_redis4,kitchen_sink3,auth_session_ratelimit2,driver_memcached2,observability,static_swagger_proxy,auth_jwt_csrf)Every affected cell shows the same shape:
h2c_sent=3597, h2c_upgraded=0, h2c_declined~2463, h2c_hang=1..4.Overall rate ≈ 1 hang per ~24,700 h2c requests, present in every run:
Related:
sse_handshake_failis also io_uring-only3 events across the same six runs, all io_uring (
observability,kitchen_sink,auth_jwt_csrf— all arm64). Possibly the same underlying accept/upgrade path; noting it here rather than opening a second issue until the cause is known.Good news
No HIGH-severity invariant has ever fired across all 288 cells:
adv_wrong_accepted,h2c_crashedandws_accepted_bad_frameare all clean. Nothing is crashing or accepting malformed input — this is a liveness/wedge issue, not a correctness one.Why it went unnoticed
Reported separately in goceleris/probatorium — the short version is that
h2c_hangis MED severity, MED only fails the gate underVALIDATE_DIFF_STRICT(never set in any workflow), and the cross-arch diff only flags a counter when it is zero on exactly one side. Because io_uring hangs on both arches, most nights it is either symmetric (invisible) or coincidentally asymmetric (reported, but non-fatal).Evidence gathered from nightly artifacts
matrix-nightly-results-{33546772656,33482366357,33243914887,33071374027,32924910721,32803245801}.