Skip to content

io_uring: connState generations pinned at 1 let a close-path ASYNC_CANCEL kill the next conn on a recycled fd #470

Description

@FumingPower3925

Summary

Across the last six nightly validation runs (288 cells, ~1.04M h2c requests), the h2c_hang Tier 1 invariant counter fires only on the io_uring engine — never on epoll or std, under an identical workload.

engine cells cells with h2c_hang rate h2c_sent total
epoll 96 0 0.0% 345,312
iouring 96 14 14.6% 345,213
std 96 0 0.0% 345,312

h2c_hang is a documented "must be zero" Tier 1 invariant (report/diff.go:invariantCounters).

Why this is a real wedge, not measurement noise

The counter has been deliberately de-noised in probatorium (validation/h2c.go):

  • h2c_intentional_rst was split out of hang so the workload's own RST-before-read no longer pollutes it — the comment notes a 3-day soak once showed h2c_hang=317K that "was almost entirely this workload noise".
  • The read budget was raised 2s → 10s specifically to eliminate ~18K false positives per soak from slow-but-correct refapps.

So a h2c_hang today means: the server accepted the connection, received a valid h2c upgrade preamble, and produced zero bytes for a full 10 seconds.

Distribution — engine-specific, not arch- or app-specific

  • by engine: {iouring: 14} — 14 of 14
  • by arch: {amd64: 7, arm64: 7} — evenly split, so not an arm64 issue
  • by refapp: spread over 7 different refapps (driver_redis 4, kitchen_sink 3, auth_session_ratelimit 2, driver_memcached 2, observability, static_swagger_proxy, auth_jwt_csrf)

Every affected cell shows the same shape: h2c_sent=3597, h2c_upgraded=0, h2c_declined~2463, h2c_hang=1..4.

Overall rate ≈ 1 hang per ~24,700 h2c requests, present in every run:

run           cells   non-zero invariant counters
33546772656     48    MED:h2c_hang=1  LOW:sse_handshake_fail=1
33482366357     48    MED:h2c_hang=2
33243914887     48    MED:h2c_hang=2  LOW:sse_handshake_fail=1
33071374027     48    MED:h2c_hang=3  LOW:sse_handshake_fail=1
32924910721     48    MED:h2c_hang=4
32803245801     48    MED:h2c_hang=2

Related: sse_handshake_fail is also io_uring-only

3 events across the same six runs, all io_uring (observability, kitchen_sink, auth_jwt_csrf — all arm64). Possibly the same underlying accept/upgrade path; noting it here rather than opening a second issue until the cause is known.

Good news

No HIGH-severity invariant has ever fired across all 288 cells: adv_wrong_accepted, h2c_crashed and ws_accepted_bad_frame are all clean. Nothing is crashing or accepting malformed input — this is a liveness/wedge issue, not a correctness one.

Why it went unnoticed

Reported separately in goceleris/probatorium — the short version is that h2c_hang is MED severity, MED only fails the gate under VALIDATE_DIFF_STRICT (never set in any workflow), and the cross-arch diff only flags a counter when it is zero on exactly one side. Because io_uring hangs on both arches, most nights it is either symmetric (invisible) or coincidentally asymmetric (reported, but non-fatal).

Evidence gathered from nightly artifacts matrix-nightly-results-{33546772656,33482366357,33243914887,33071374027,32924910721,32803245801}.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/engineEngine interface or implementationbugSomething isn't workingphase/1-enginesPhase 1: Engine implementationsplatform/linuxLinux-specific (io_uring, epoll)priority/parallelCan be worked in parallelsize/L~3-5 days of work

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions