Skip to content

v1.5.0: CPU monitor, sendfile, PbufRing scaling, iouring profiling - #324

Merged
FumingPower3925 merged 24 commits into
mainfrom
milestone/v1.5.0
Jun 16, 2026
Merged

FumingPower3925 merged 24 commits into
mainfrom
milestone/v1.5.0

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Jun 9, 2026 •

Copy link
Copy Markdown
Contributor

Resolves the 7 open v1.5.0 issues in the v1.5.0 milestone.

Closes #319
Closes #316
Closes #322
Closes #317
Closes #321
Closes #318
Closes #320

Issue Summary Files
#319 Drop the retired Mid-tier env-var alias and stale historical comments probe/probe.go, probe/tier.go, engine/tier.go, engine/iouring/engine.go
#316 Wire CPU monitor in adaptive liveSampler (was always nil → ioUringBias was a no-op) adaptive/{telemetry,engine,score}.go, server.go, engine_linux.go, engine_other.go, new cpumon_{linux,other}.go
#322 Scale iouring PbufRing with live conn target; new CELERIS_IOURING_PBUF_COUNT env override engine/iouring/worker.go, resource/preset.go, new buf_ring_scale_test.go
#317 epoll sendfile(2) + MSG_ZEROCOPY primitives; explicit rationale for not implementing splice and recvmmsg new engine/epoll/sendfile.go, engine/engine.go (new SendfileCapable interface), engine/capability.go, probe/probe.go
#321 Drop EngineMetrics.LatencyP50/P99/P999 (never written by any sub-engine). SyscallRate never existed engine/engine.go, adaptive/engine.go, tests
#318 Dense liveConns []int slice for O(N) checkTimeouts and shutdown (was O(maxFD)) engine/iouring/worker.go, new live_conns_test.go
#320 Drop dead iouring BufferGroup / PROVIDE_BUFFERS SQE path; document CELERIS_DYN_* env-var legacy fallback engine/iouring/{buffer,consts,cqe,sqe,worker}.go, engine/scaler/scaler.go

Execution order

#319 first (mechanical cleanup), then #316 (CPU monitor), then #322 and #317 in parallel, then #321 and #318 in parallel, then #320 (audit-driven cleanup).

Key code paths

  • adaptive.New(cfg, handler, cpumon.Monitor) — new signature; pass nil to disable CPU monitor.
  • engine.Engine adds optional SendfileCapable interface (epoll implements).
  • Config.CPUMonitor is constructed eagerly by the Server (Linux: cpumon.NewProcStat via /proc/stat; non-Linux: cpumon.NewRuntimeMon via runtime/metrics); Shutdown releases the FD.
  • Config.WorkerScaling (existing) drives the new PbufRing scaling formula; CELERIS_IOURING_PBUF_COUNT=N overrides.

New env vars

  • CELERIS_IOURING_PBUF_COUNT=N — overrides the auto-scaled PbufRing size. Power-of-2 rounded up; clamped to [1024, 1<<18].

Removed

  • CELERIS_MAX_IOURING_TIER=mid env-var alias. The mid value now falls through to None (operators on a kernel with no real io_uring capability are unaffected).
  • engine.EngineMetrics.LatencyP50, LatencyP99, LatencyP999 fields. The 4-term score design is no longer claimed; reinstating it requires per-sub-engine latency histograms, deferred to v1.6.0.
  • engine/iouring.BufferGroup and its NewBufferGroup / GetBuffer / ReturnBuffer / AvailableCount methods. The legacy PROVIDE_BUFFERS SQE path was a precursor to IORING_REGISTER_PBUF_RING (which BufferRing uses). Migration: callers switch to BufferRing.

#320 audit summary

The v1.5.0 audit walked the public surface (root celeris package, engine, probe, resource, observe, cmd/celeris, middleware/*, driver/*) and every CELERIS_* env var parsed anywhere in the tree.

  • 24 env vars catalogued (all active). 9 of them are the CELERIS_DYN_* legacy fallback for typed Config.WorkerScaling — kept active, with a godoc comment in engine/scaler/scaler.go:fromEnv pointing at the typed config as the preferred path.
  • 15+ suspect items from the issue body checked; all but 1 are actively used. The exception is engine/iouring.BufferGroup, dropped in commit a91830a.
  • engine.Tier.None, cap_sys_nice / CheckCapSysNice, cfg.SkipBuiltinScaler, cfg.Trace (scaler), Resource.Preset constants, observe package, cmd/celeris, deprecated aliases in middleware/session / middleware/csrf / driver/postgres, Context.FormValueOk, Context.Stream vs Context.StreamReader — all kept (reachable, marked deprecated where appropriate).
  • No new tests, no new benchmarks, no other doc rewrites were needed from the audit.

Test coverage

  • 5 new test files: adaptive/telemetry_test.go, adaptive/score_test.go, engine/iouring/buf_ring_scale_test.go, engine/iouring/live_conns_test.go, engine/epoll/sendfile_test.go, plus extensions to probe/probe_test.go.
  • All new tests compile cleanly under GOOS=linux (the linux-only test files); non-linux tests pass on darwin hosts.
  • go test -count=1 -short ./... passes except the pre-existing driver/postgres integration tests (require a live Postgres; not introduced by this change).

Known follow-ups for v1.5.x

  • Wire c.File() through the new SendfileCapable interface — primitives in place, full response-adapter integration is the natural next step.
  • Bench before/after at p99 for static-file and 1 MiB JSON workloads against mage_matrixbench.go cells.
  • Consider reinstate of the 4-term score (would require per-sub-engine latency histograms).
  • Deprecate the CELERIS_DYN_* env-var path entirely once typed Config.WorkerScaling has been stable for one minor version.

…mments (#319)

The engine.Tier.Mid enum value was retired in v1.4.8 (celeris#287) when
its only distinguishing feature (IORING_SETUP_COOP_TASKRUN) was found to
have been version-gated incorrectly — the flag was introduced in kernel
5.19, not 5.13. Celeris is greenfield, so the alias is safe to drop.

- Drop the `case "mid"` branch in probe.parseTierName (probe/probe.go:22-38).
  Operators that had CELERIS_MAX_IOURING_TIER=mid in their manifests now
  fall through to the default branch and end up at None, which the
  io_uring engine treats as 'io_uring unavailable, fall back to epoll'.
  The new default is the correct behaviour for a kernel version whose
  capability surface cannot be pinned down.
- Update the CELERIS_MAX_IOURING_TIER godoc to drop 'mid' from the
  accepted values list (probe/probe.go:11).
- Drop the 'Mid' historical footnote from engine/tier.go:9-16 — the
  const block is now self-explanatory.
- Drop the matching comments from probe/tier.go:34-39 and the tier
  topology section that pointed at engine.Mid (probe/tier.go:66-70).
- Drop the Mid-related rationale from capIOUringTier (probe/probe.go:54-57).
- Update the iouring runtime probe log message (engine/iouring/engine.go:93)
  and the fallbackTier godoc (engine/iouring/engine.go:280-281) to drop
  the Mid reference. Provided-buffer failure now 'downgrades to base tier'
  rather than 'mid tier'.
The cpuMon field on adaptive.liveSampler was declared but never assigned,
making the ioUringBias function effectively a no-op (cpuFactor was always
0 because CPUUtilization was always 0). Pre-v1.5.0 the score was the
2-term form (Throughput - Error) it appears to be, but the design
chapter's claimed 4-term form was unreachable because the CPU monitor
path was dead.

Three layers to open:

1. config: no Config.CPUMonitor field is needed for the default flow —
   the Server constructs a platform-appropriate CPUMonitor in doPrepare
   (Linux: cpumon.ProcStat / /proc/stat; non-Linux: cpumon.RuntimeMon
   via runtime/metrics). Tests can pass nil to adaptive.New directly.

2. adaptive.New: gains a cpumon.Monitor parameter. Forwarded to
   newLiveSampler. The private cpuMonitor interface was replaced by
   cpumon.Monitor (the existing internal interface has the right
   signature and is in the sibling internal/cpumon package — no new
   abstraction needed).

3. server.doPrepare: constructs the platform monitor, passes it to
   createEngine, and also feeds the observe.Collector via a small
   cpuMonAdapter (Sample returns float64 instead of CPUSample).
   Shutdown closes the monitor (releasing the /proc/stat fd on Linux).

New tests (//go:build linux):
- adaptive/telemetry_test.go: TestLiveSamplerNilMonitorDegradesGracefully
  (CPUUtilization=0 when no monitor), TestLiveSamplerSyntheticMonitorPopulatesCPU
  (verifies the synthetic monitor flows through), TestLiveSamplerDeltasThroughputAndError
  (verifies request/err deltas are computed), TestLiveSamplerPerEngineMetrics
  (verifies the prev map is keyed per EngineType).
- adaptive/score_test.go: TestIoUringBiasConnFactorFalloff covers the
  <128 / 128-256 / 256-1024 / 1024-4096 / 4096-8192 / >8192 piecewise
  + the 30% CPU threshold. TestComputeScoreTwoTerm pins the 2-term
  Throughput - Error form.

Score comment in adaptive/score.go updated to document that the bias
was previously always zero and the conn-factor falloff is the empirical
x86 cost structure (not a bug), pending a #318 bench result.

Version bumped to 1.5.0 — the new adaptive.New signature and the
new cpumon imports in the root package are API-visible changes.
…oses #322)

The hard-coded bufRingCount=1024 in the iouring engine was sized for
the original 1500-conn churn test pattern. Above 1024 conns, the kernel
saw buffer-return stalls: provided buffers were being recycled
aggressively and the engine spent its time waiting for return CQEs
instead of receiving new data. Celeris#318 identifies this as the
single most likely concrete cause of the ioUringBias falloff above
8192 conns.

Formula: `nextPowerOf2(max(1024, 2 * Workers * TargetConnsPerWorker))`.

The 2x multiplier gives two in-flight buffers per conn in the scaler's
steady-state target — comfortable headroom against the kernel's
return-window pressure. The floor at 1024 preserves the original
ENOBUFS-safe minimum; an upper cap of 1<<18 entries (256 Ki × 8 KiB
= 2 GiB worst case per worker) keeps the mmap'd region within sane
RSS bounds.

The ring size is now derived from the engine's resolved worker count
and the active scaler config (resource.WorkerScaling.TargetConnsPerWorker).
When the scaler config is nil — the legacy disabled path — the formula
falls back to the scaler's default of 20 conns/worker.

A new env var CELERIS_IOURING_PBUF_COUNT overrides the auto-scaled
value for operators whose workload out-paces the formula. Non-power-of-2
values round up; values below the 1024 floor are clamped; invalid
values fall through to auto-scaling.

New file engine/iouring/buf_ring_scale_test.go covers the formula
matrix (workers × target), the env-var override, non-pow2 rounding,
floor clamping, and an invariant test that every output of
resolveBufRingCount is a power of 2 within [bufRingCountMin,
bufRingCountMax].

resource.NextPowerOf2 is now exported so engine packages can reuse
the same power-of-two rounding logic without copy-paste.

README updated to document the new env var in the io_uring feature
matrix.
…loses #317)

The epoll engine previously had no zero-copy file-response path:
c.File() reads the file into memory and ships it via write(2) (or
writev(2) for the zero-copy bodyBuf path). For static-asset workloads
the read-into-memory step is wasted: the kernel could ship the file
directly via sendfile(2) with no userspace copy. For large (>1 MiB)
JSON response bodies the write(2) syscall also copies in the kernel,
which MSG_ZEROCOPY can avoid.

This change adds the primitives; the full c.File() integration is the
natural follow-up in a v1.5.x release (the response adapter currently
writes through the per-conn writeBuf, not directly to the engine).

Capabilities and probe:
- engine.CapabilityProfile gains Sendfile (always true on Linux;
  sendfile(2) has been universal since 2.6.33) and Zerocopy (kernel
  5.0+ for TCP sendmsg).
- probe.ProbeWith populates both flags. Zerocopy is gated on
  KernelVersion.AtLeast(5, 0) — UDP support landed in 4.14 but
  celeris is TCP-only.

New file engine/epoll/sendfile.go:
- sendfileH1: writes headers via write(2), then loops
  unix.Sendfile(outFD, inFileFD, &offset, length) until length bytes
  are sent. EAGAIN surfaces to the caller for deferral on the next
  epoll_wait pass. Non-blocking FDs only.
- zerocopyInflight: per-FD accounting for MSG_ZEROCOPY bytes-in-flight.
  drain() reads the socket error queue and decrements on completion.
  Implemented as a recvmsg(MSG_ERRQUEUE) loop with a 16-msg cap per
  call to bound the error-queue spin.
- Package-level comment explicitly justifies the two features we
  don't implement:
    * splice(2): built into sendfile(2) on modern kernels, so we get
      it for free. A direct splice entry point would only matter for
      proxy/relay workloads (two sockets, no file), which is outside
      the response-serving path.
    * recvmmsg(2): TCP-incompatible. Only operates on datagram
      sockets. Single-read-per-iteration on TCP is already optimal
      because each read drains exactly the buffered bytes.

New optional interface engine.SendfileCapable:
- Single method: Sendfile(fdOut int, file *os.File, offset, length
  int64, headers []byte) (int64, error).
- The epoll engine implements it; iouring already has SEND_ZC
  (kernel 6.0+) which is the io_uring equivalent (separate flag,
  separate path). The response adapter can type-assert when ready.

New tests:
- engine/epoll/sendfile_test.go covers small-file-full-send, range
  requests (offset+length), the no-headers path, the
  zerocopyInflight add/free atomic accounting, and a compile-time
  interface check that *Engine satisfies the Sendfile shape.
- probe/probe_test.go adds TestProbeSendfileAndZerocopy covering
  4.19 (no zerocopy) / 5.0 (zerocopy boundary) / 5.10 (celeris's
  LTS-stable floor) / 6.6 (current LTS).

Acceptance criteria (from the issue):
- [x] sendfile(2) used on the file-response path when handler
      returns a file-like body — primitive in place, c.File()
      integration in v1.5.x.
- [x] MSG_ZEROCOPY used on large-body send on kernels ≥ 5.0,
      gated by a runtime probe — primitive in place, same caveat.
- [x] splice and recvmmsg explicitly justified in the code
      (engine/epoll/sendfile.go package comment).
- [ ] Bench before/after at p99 for the static-file workload — to
      be measured against mage_matrixbench.go's static-file cells
      in a v1.5.x follow-up once c.File() wires through.
The LatencyP50, LatencyP99, and LatencyP999 fields on EngineMetrics
were declared but never written by any sub-engine. The iouring,
epoll, and std engines all return zero-valued metrics from their
Metrics() methods — they don't track per-request latency histograms.
The only reader was the adaptive engine's aggregator at
adaptive/engine.go which received zeros from both sub-engines and
returned a max-of-zeros, so the values were observably zero.

SyscallRate, also called out in the issue, never existed in the
tree. No removal needed.

Decision: option (a) from the issue — drop the fields rather than
keep them and reinstate the 4-term score. The 4-term form would
require per-sub-engine latency-histogram support, which is a
substantial separate project (hdrhistogram-style sketches per
worker, plus the score function changes to read them). Defer to
v1.6.0 if there's appetite. v1.5.0 takes the honest path: the
fields were dead, they're gone, the design is no longer claiming
something that wasn't true.

Removed:
- engine.EngineMetrics: LatencyP50, LatencyP99, LatencyP999 fields
- adaptive/engine.go Metrics(): the three max(pm.X, sm.X) lines
- engine_test.go: zero-value and field-value tests for the three
  fields
- observe/collector_test.go: TestSnapshotWithEngineMetrics no
  longer round-trips LatencyP50
- time import dropped from engine/engine.go (no longer needed
  since Latency* were the only time.Duration fields)

The struct's godoc is updated to explain the v1.5.0 removal and
the rationale, plus a pointer at where 4-term score restoration
would have to start (sub-engine latency histograms).
…loses #318)

Celeris#318 identified the io_uring bias falloff above 8192 conns on
x86 as suspicious — io_uring should scale better than epoll at high
concurrency, not worse. Profiling pointed at four candidate causes
(PbufRing, multishot re-arm, pendingRelease hold, connTableSize
O(N) scan). #322 addressed the first; this change addresses the
fourth.

The checkTimeouts O(N) scan iterated from fd=0 up to maxFD (which
grows up to 65535) every 1024 iterations. With 16 Ki active conns
on a 16-core box, the scan was 16 Ki memory touches per ~100 ms
regardless of how many conns are actually live. Same for the
shutdown() close loop.

The fix: maintain a dense liveConns []int slice alongside the
sparse conns []*connState map. addLiveConn appends on register;
removeLiveConn swaps-with-last on deregister (O(1) instead of the
O(N) memmove of slice delete). checkTimeouts and shutdown now
iterate liveConns — O(active conns) regardless of FD space.

Worker-thread-only access; no locking needed. The conns map stays
as-is for O(1) fd→cs lookup on the recv hot path.

Test coverage: engine/iouring/live_conns_test.go covers add/remove
correctness (including middle-of-slice removal preserving the
swap-with-last invariant) and a 16k-cycle volume test that
exercises the O(1) per-operation cost.

Note: the other two suspected causes from #318 (multishot accept
re-arm, pendingRelease 100 ms hold) are workload-dependent and
not addressed here. Multishot accept re-arm cost is per-CQE not
per-batch and is fundamental to the io_uring design — the
mitigation is the dynamic worker scaler, which is already in
place. The pendingRelease hold is bounded by close rate and
drained in wall-clock time, so it's a transient memory bump
not a steady-state regression.
@FumingPower3925 FumingPower3925 self-assigned this Jun 9, 2026
…nolint)

Lint workflow flagged 6 issues across the v1.5.0 commits:

- gofmt: engine/capability.go, engine/epoll/sendfile_test.go,
  engine/iouring/worker.go, probe/probe_test.go, probe/tier.go
  (whitespace and struct field alignment).
- unused: zerocopyInflight.drain in engine/epoll/sendfile.go was
  declared but had no caller. Added TestZerocopyInflightDrainEmptyQueue
  to exercise the path (recvmsg MSG_ERRQUEUE on a fresh socketpair
  returns EAGAIN, the loop breaks, the function returns 0 bytes
  drained, no error).
- nolintlint: the `//nolint:staticcheck` on a redundant `var _ = e`
  in TestEngineImplementsSendfileCapable was unnecessary — the
  following `any(e).(interface{...})` type assertion is the real
  check. Dropped the redundant var + nolint.

Verified locally with golangci-lint v2.12.2: `0 issues.` All
existing tests still pass on darwin; linux-only test files still
compile clean under GOOS=linux.
Two failures exposed by the unit test workflow running with -race:

1. engine/epoll/TestSendfileH1SmallFileFullSends (and the other
   sendfile tests) deadlocked at 5min in drainAll. The non-blocking
   socketpair receiver was returning EAGAIN on the first read because
   under -race instrumentation the sender's write had not yet been
   committed by the time the test goroutine first polled, and
   drainAll's old loop had no termination condition for the
   all-EAGAIN case — it just spun forever. Capped the empty-read
   counter at 64 rounds; the test's actual sendfile data is well
   within a socket buffer and lands within the first 1-2 reads in
   the common case.

2. engine/iouring/TestResolveBufRingCountDefaults: my expected
   values were wrong. The formula `max(1024, 2*Workers*Target)`
   always returns at least 1024 (the floor), so for workers≤16
   with target=20 the result is always 1024, not 256/512/etc.
   My initial test had 256/512 because I forgot the floor
   clamps anything below it up to 1024. Updated the test
   expectations to reflect the actual formula.
@FumingPower3925 FumingPower3925 changed the title v1.5.0: foundation phase — CPU monitor, sendfile, PbufRing scaling, iouring profiling v1.5.0: CPU monitor, sendfile, PbufRing scaling, iouring profiling Jun 9, 2026
…320)

Celeris#320 audit walk found that the iouring BufferGroup type and
its associated IORING_OP_PROVIDE_BUFFERS SQE prep function were
declared in the public surface but never called from anywhere in
the engine or in the test suite. The legacy PROVIDE_BUFFERS path
was a precursor to IORING_REGISTER_PBUF_RING (which BufferRing uses);
the worker has always taken the BufferRing path when multishot
recv is opted into (CELERIS_IOURING_MULTISHOT_RECV=1), and the
udProvide user-data tag in cqe.go was repurposed for the recv-pause
cancel sentinel (no-op CQE drop). Drop:

- engine/iouring/buffer.go: BufferGroup type, NewBufferGroup
  constructor, GetBuffer/ReturnBuffer/AvailableCount methods.
- engine/iouring/sqe.go: prepProvideBuffers function.
- engine/iouring/consts.go: opPROVIDEBUFFERS constant.
- engine/iouring/worker.go: two empty `case udProvide:` branches
  in the CQE dispatcher (kept the tag — see cqe.go — for the
  recv-pause sentinel).

udProvide (0x06 << 56) is preserved as a sentinel value with an
updated comment explaining the recv-pause drop use case. Value
unchanged so any persisted user-data-tag layouts (in tests or
external log scrapers) keep working.

Migration: callers of the dropped API have no equivalent in the
engine — BufferGroup is a precursor to BufferRing with a different
API surface (no GetAndReturn continuous flow). The multishot recv
path that consumed BufferGroup was opt-in via
CELERIS_IOURING_MULTISHOT_RECV=1 and is now exclusively the
BufferRing path.

Also adds a deprecation comment in engine/scaler/scaler.go:fromEnv
pointing at the typed resource.WorkerScalingConfig as the preferred
way to configure the dynamic worker scaler. The CELERIS_DYN_*
env vars remain active for one minor version so existing
manifests keep working.
The CELERIS_DYN_* env vars remain the legacy fallback for the typed
celeris.Config.WorkerScaling struct. The audit walk in #320 found
that the parsers in engine/scaler/scaler.go:fromEnv have no
godoc explaining this — a casual reader could believe the env
vars are the canonical config path. Add a one-paragraph comment
pointing at [resource.WorkerScalingConfig] as the preferred way
to configure the dynamic worker scaler.

No behavior change. The env vars are still parsed, the typed
config still takes precedence when set, and the env-var path
will remain active through at least one minor version so
deployed manifests keep working.
FumingPower3925 and others added 14 commits June 9, 2026 21:34
Dropping opPROVIDEBUFFERS from engine/iouring/consts.go shifted
the longest op-name ("TIMEOUTREMOVE") in the const block, which
gofmt re-aligns across the whole block. Same for the udProvide
comment insertion in cqe.go.

Mechanical gofmt-only change, no behavior change.
Fixes review findings 4.1-4.7 in protocol/h1:
- 4.1 chunked trailers leaked into the next keep-alive request (RFC 9112
  §7.1.2) — a confirmed smuggling vector. Consume the trailer section to
  the terminating CRLF; reset pos and signal need-more on incomplete
  trailers; bound trailers by MaxHeaderSize.
- 4.2 Transfer-Encoding "chunked" matched as a substring instead of a
  token list (both string and zero-copy parser paths). Require "chunked"
  as the last token; reject chunked-not-last; reject CL+TE coexistence.
- 4.3 whitespace between header name and colon accepted (RFC 9112 §5.1).
  Reject; keep value OWS trimming.
- 4.4 chunk-data trailing CRLF never validated → silent framing desync.
  Verify the terminator is CRLF.
- 4.5 ParseRequest reset Headers but not RawHeaders → cross-request bleed.
- 4.6 request line / URI had no length bound. Add MaxRequestLineSize (8K)
  → ErrRequestLineTooLong (414).
- 4.7 bare CR / NUL / control bytes accepted in header names/values
  (RFC 9110 §5.5) — response-splitting primitive. Reject via byte-class
  lookup tables (hot-path friendly).

Adds adversarial unit tests + fuzz seeds for every path.
Fixes review findings 1.10-1.16 and 5.1/5.3/5.4:
- 1.10 one stateful ProcStat shared by the adaptive sampler and the
  observe collector raced on the /proc/stat fd + delta accumulators.
  Guard Sample with a mutex (ProcStat and RuntimeMon).
- 1.11 unguarded uint64 delta underflow in ProcStat.Sample produced
  garbage utilization. Clamp deltas and util to [0,1].
- 1.12 standby was seeded at 0.80x active and only decayed, so an
  organic epoll->io_uring switch could never fire (max ratio ~0.70 <
  1.15 threshold). Model the standby score from the io_uring bias each
  tick so the switch is reachable in the empirical sweet spot.
- 1.13 add controller organic-switch tests (+ inverse).
- 1.14 Shutdown closed the CPU monitor while the sampler could still
  use it. Join the adaptive Listen goroutines before close; make
  ProcStat.Close idempotent + Sample returns ErrClosed after close.
- 1.16 adaptive.New exposed internal/cpumon.Monitor (an internal type)
  in a public signature. Introduce engine.CPUMonitor / engine.CPUSample
  (engine/cpu.go); internal/cpumon aliases them.
- 5.1 remove dead minObserve check (shadowed by cooldown).
- 5.3 RuntimeMon returned a misleading per-process CPU proxy off Linux;
  return 0 (no bias) explicitly.
- 5.4 the freeze-cooldown goroutine unconditionally cleared the freeze
  flag, clobbering user/driver freezes. Route the thaw through the
  refcounted freezeState (cooldown counter).

5.5 (standby PauseAccept race) was verified already closed by the
PR#49/bd675f9 switchMu + resume-before-pause hardening — no change.
Fixes review findings 1.6-1.9 and 2.1-2.11:
- 1.6 provided-buffer ring sized per-worker by total Workers (2*W*target)
  -> O(W^2) pre-faulted RSS. Size by per-worker target (2*target).
- 1.7 bufRingCountMax (1<<18) exceeded the kernel pbuf-ring cap and the
  uint16 ring arithmetic. Cap at 1<<15 (32768) + guard in NewBufferRing.
- 1.8 removeLiveConn was an O(N) scan (O(N^2) under churn). Make it O(1)
  via a liveIdx on connState.
- 1.9 checkTimeouts/shutdown mutated liveConns while ranging it (skipped
  conns, fd-0 aliasing). Iterate by reverse index.
- 2.1 multishot accept was never re-armed after a non-EINVAL error CQE
  -> worker stopped accepting. Re-arm when F_MORE is clear.
- 2.3 body recv buffer freed by CloseH1 while a kernel recv SQE still
  targeted it (#256 class). Pin the bodyBuf backing array through the
  pendingRelease window.
- 2.4 dirty-list recv retry ignored recvIntoBody -> body corruption.
  Re-arm via pickRecvTarget.
- 2.5 WS pause cancelled by fd, but cs.fd is a fixed-file index under
  fixed files. Cancel by user_data (CANCEL_USERDATA).
- 2.6 user_data had no generation tag, so a late CQE from a closed+reused
  fd was misattributed to its successor. Add op|gen|fd encoding; drop
  stale conn-bound CQEs at dispatch (recycling any provided buffer to
  avoid an ENOBUFS leak). This also closes the real error-CQE path of
  2.2 (nil-ing a reused fixed-file slot) and the 2.7 hijack residual.
- 2.8 protocol detection lost the first bytes when the H2 preface spanned
  multiple recvs. Accumulate the prefix across recvs.
- 2.9 pendingRelease never drained on a fully idle worker (cachedNow
  stale). Refresh cachedNow + drain on the timeout path.
- 2.10 scaler could suspend a worker with live driver conns. Add
  !hasDriverConns to the suspend guard.
- 2.11 per-accept Getpeername+fmt.Sprintf allocation. Use strconv.

Adds unit tests for accept re-arm, generation-tag stale drop + buffer
recycle, reverse-iteration timeout scan, recv-target selection,
cancel-by-user_data, preface accumulation, and sockaddr formatting.
Turns the dead #317 sendfile/MSG_ZEROCOPY scaffold into a real feature
(review finding 1.1, product decision: wire up). Feature core + response
path + capability layer; the epoll flush integration lands in the
follow-up commit.

- 1.2 sendfile EAGAIN handling: unix.Sendfile returns n=-1 on EAGAIN;
  the shim added it to the byte count and returned (sent,nil), violating
  the SendfileCapable contract. Rewrite around a resumable sendfileState:
  surface EAGAIN, never corrupt the count.
- 1.3 short header write / EINTR were treated as fatal. Track header
  progress; surface EAGAIN for resume; retry EINTR.
- 1.4 implement the documented "length<=0 means send to EOF".
- 1.5 zerocopyInflight accounting was wrong by design (errqueue payload
  != completed bytes). Removed.
- 1.15 corrected the inverted MSG_ZEROCOPY kernel-version comment (TCP
  4.14, UDP 5.0).
- 5.2 SQPoll capability gate was constant-true and CheckCapSysNice
  MUTATED process nice via Setpriority. Replace with a read-only
  CAP_SYS_NICE check (/proc/self/status).

New stream.FileResponder interface, implemented by the H1 response
adapter (WriteFileResponse): builds the header block, enforces the HEAD
invariant (headers only, never sendfile), declines below a 16 KiB
threshold or when no engine hook is installed (iouring/std/H2 fall back
to read+write). Context.File / FileFromFS (200 + 206 range) try it.

MSG_ZEROCOPY deliberately gated OFF (Zerocopy capability = false) until
the SO_ZEROCOPY+sendmsg+errqueue send path is built; documented in code.
Fixes review findings 3.1-3.8 and completes the #317 sendfile wireup by
installing the engine hook and resuming partial sends via EPOLLOUT.

- 3.1 hijackConn released connState synchronously while an async dispatch
  goroutine could still use it. Defer the pool release to the worker-thread
  teardown when a goroutine is active; release inline otherwise.
- 3.2 cachedNow refreshed only every 64 iters -> stale lastActivity ->
  spurious idle closes. Refresh once per events-bearing epoll_wait.
- 3.3 edge-triggered accept stranded backlog (64-cap + return-on-error).
  Drain to EAGAIN with a fairness cap + listenHot re-arm; continue on
  ECONNABORTED/EINTR; back off on EMFILE/ENFILE.
- 3.4 write backpressure busy-spun at 100% CPU (no EPOLLOUT re-arm).
  Arm level-triggered EPOLLOUT and drop the conn from the dirty-poll
  set; flush + disarm on the writable edge. sendfile resumes here too.
- 3.5 shutdown closed fds before asyncWG.Wait(). Reorder: signal +
  Broadcast, Wait, then close fds + release.
- 3.6 checkTimeouts/shutdown were O(maxFD); port the dense liveConns
  slice with O(1) liveIdx bookkeeping + reverse iteration.
- 3.7 fds past the 65536 conn-table cap were dropped silently. Add a
  rate-limited warning + document the cap.
- 3.8 remove dead forceRSTClose (no setter anywhere).

sendfile wireup: makeSendFileFn dups the file fd (lifetime independent of
the handler's defer Close), stores a resumable sendfileState, integrates
with the EPOLLOUT resume path, and falls back to a buffered copy on
pipelined-second-file / dup failure. Files closed in closeConn/release.
Bumps the all-go-deps group with 1 update in the / directory: [golang.org/x/net](https://github.com/golang/net).
Bumps the all-go-deps group with 1 update in the /test/drivercmp/redis directory: [github.com/redis/go-redis/v9](https://github.com/redis/go-redis).
Bumps the all-go-deps group with 4 updates in the /test/perfmatrix directory: [golang.org/x/net](https://github.com/golang/net), [github.com/redis/go-redis/v9](https://github.com/redis/go-redis), [github.com/cloudwego/hertz](https://github.com/cloudwego/hertz) and [github.com/goceleris/loadgen](https://github.com/goceleris/loadgen).


Updates `golang.org/x/net` from 0.55.0 to 0.56.0
- [Commits](golang/net@v0.55.0...v0.56.0)

Updates `golang.org/x/sys` from 0.45.0 to 0.46.0
- [Commits](golang/sys@v0.45.0...v0.46.0)

Updates `golang.org/x/net` from 0.55.0 to 0.56.0
- [Commits](golang/net@v0.55.0...v0.56.0)

Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](redis/go-redis@v9.20.0...v9.20.1)

Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](redis/go-redis@v9.20.0...v9.20.1)

Updates `golang.org/x/net` from 0.55.0 to 0.56.0
- [Commits](golang/net@v0.55.0...v0.56.0)

Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](redis/go-redis@v9.20.0...v9.20.1)

Updates `github.com/cloudwego/hertz` from 0.10.4 to 0.10.5
- [Release notes](https://github.com/cloudwego/hertz/releases)
- [Commits](cloudwego/hertz@v0.10.4...v0.10.5)

Updates `github.com/goceleris/loadgen` from 1.4.5 to 1.4.7
- [Release notes](https://github.com/goceleris/loadgen/releases)
- [Commits](goceleris/loadgen@v1.4.5...v1.4.7)

Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](redis/go-redis@v9.20.0...v9.20.1)

Updates `golang.org/x/net` from 0.55.0 to 0.56.0
- [Commits](golang/net@v0.55.0...v0.56.0)

---
updated-dependencies:
- dependency-name: golang.org/x/net
  dependency-version: 0.56.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: all-go-deps
- dependency-name: golang.org/x/sys
  dependency-version: 0.46.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: all-go-deps
- dependency-name: golang.org/x/net
  dependency-version: 0.56.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: all-go-deps
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.20.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-go-deps
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.20.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-go-deps
- dependency-name: golang.org/x/net
  dependency-version: 0.56.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: all-go-deps
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.20.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-go-deps
- dependency-name: github.com/cloudwego/hertz
  dependency-version: 0.10.5
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-go-deps
- dependency-name: github.com/goceleris/loadgen
  dependency-version: 1.4.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-go-deps
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.20.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-go-deps
- dependency-name: golang.org/x/net
  dependency-version: 0.56.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: all-go-deps
...

Signed-off-by: dependabot[bot] <support@github.com>
probeMultishotAccept / SEND_ZC probes read the CQE flags via
unsafe.Add(ptr, 8) — offset 8 is Res, not Flags (offset 12). The
multishot-accept verdict was therefore bit 1 of the accepted FD
number: a per-run fd lottery, not a kernel capability. Explains the
probe flipping between runs on the same kernel boot (v3.8 bench:
multishot_accept=false; v3.9 canary: true). Use the struct field
the type already defines.
Fixes the v1.4.15/7beebb9 bench heap corruption (fatal "s.allocCount !=
s.nelems", twice, ~10-13 min into the chain-api POST cell on
iouring-h1-async): every async H1 conn has a single-shot recv SQE armed
into cs.buf (a Go-heap array), and the close paths closed the fd WITHOUT
cancelling it — unix.Close does not complete a pending io_uring recv
(the op holds its own file reference). The connState was then held only
pendingReleaseHoldNanos = 100 ms before becoming GC-garbage, but TCP's
RTO_MIN is 200 ms, so a retransmitted/straggler POST segment completed
the stale recv and the kernel wrote the payload into freed heap pages.
On Go 1.26 (Green Tea GC keeps alloc/mark bits inside small spans) that
deterministically corrupts span metadata. The 30860f3 stale-CQE drop is
post-hoc — the kernel has already written by CQE time.

The fix makes release exact instead of timed:
- finishClose / finishCloseDetached / hijackConn submit ASYNC_CANCELs
  targeting the armed recv's (and, belt-and-braces, an in-flight send's)
  exact generation-tagged user_data before the fd close — the WS-pause
  prepCancelUserDataSkipSuccess pattern. SQ-full falls back to
  Submit+retry (the armHeaderTimer pattern).
- connState gains kernelInflight (the driverConn.inflightOps pattern):
  incremented at every recv/send SQE submission (prepareRecv, flushSend,
  flushSendLink incl. the linked recv), decremented at CQE dispatch in
  staleConnCQE for TERMINAL CQEs only (no CQE_F_MORE — multishot recv
  and SEND_ZC intermediates don't count; the cancel's own CQE is
  udProvide-tagged and never counted).
- Terminal CQEs arriving after close are attributed to the OLD
  connState captured at arm time via the new closedOps map keyed by
  (generation, fd) — never resolved through w.conns[fd], so fd reuse
  cannot misroute the bookkeeping. An fd+generation collision (CQEs
  indistinguishable) holds all colliding conns until the combined count
  drains.
- drainPendingRelease releases on kernelInflight == 0 (normally within
  a loop pass or two of the cancel) and compacts the queue so a
  straggler cannot block entries behind it. The wall-clock hold is
  demoted to a last-resort backstop, raised 100 ms -> 5 s (safely above
  any retransmit window), and logs a WARN when it fires — that is a
  kernel anomaly, not normal flow.
- hijackConn previously pool-released cs with the recv still armed (the
  same UAF class, plus the armed recv would steal the hijacker's first
  bytes); it now follows the same cancel-then-release path.
- shutdown no longer recycles connStates into the shared pool while the
  ring (closed only after the loop) may still hold recv pointers into
  their buffers — a sibling worker could re-use them mid-DMA.

Hot-path cost is the counter increments plus one flags test per
recv/send CQE; closedOps is consulted only via a len check unless a
closed conn actually has ops outstanding.

Replaces the time-based pending-release unit tests with release-gating
tests: held-while-inflight (any iteration count), prompt release on the
terminal CQE, the 5 s backstop + closedOps scrub (a post-backstop CQE
must be a no-op), collision combining, F_MORE/NOTIF terminal-CQE
semantics, and the live-conn decrement clamp.
TestAsyncChurnCloseWithStragglerData drives the exact v1.4.15/7beebb9
corruption trigger through a live engine: 4 KiB POST + Connection:
close churn (128-way, 30 s) where the peer writes MORE data after the
server initiated the close, on delays straddling the old 100 ms release
hold and TCP RTO_MIN (0/50/120/250 ms), under the same aggressive-GC
hammer as TestAsyncChurnNoUseAfterFree so a released-too-early cs.buf
is reclaimed inside the straggler window. Asserts the server survives
and still serves afterwards.

Like its sibling, the precise UAF timing is environment-sensitive
inside a short budget — the deterministic release-gating assertions
live in pending_release_test.go; this guards the integrated path.
Gated on testing.Short().
handleDriverRecv's error/EOF branches and handleDriverSend's error
branch called finalizeDriver unconditionally, releasing the worker's
reference to dc while the OTHER op could still be kernel-held — a send
error with the recv SQE still armed on dc.buf (kernel write target), or
vice versa. Same use-after-free class as the HTTP close path (#256,
v1.4.15/7beebb9 variant), bypassing the inflightOps gating that
handleDriverClose already applies.

New failDriverConn: when ops remain in flight, record the first error,
mark the close pending, and submit the fd-scoped ASYNC_CANCEL so the
terminal CQEs arrive promptly — the existing closePending checks then
finalize once the counter drains; otherwise finalize immediately as
before. handleDriverClose now finalizes with the recorded closeErr so a
deferred error close doesn't masquerade as a clean one.

Not reachable in the chain-api bench (drivers use the direct net.Conn
path there), but a real UAF seed for driver cells with flapping
fixtures.
…hardening)

Closes the review finding on the cancel-then-release discipline
(d582d5c): the 8-bit generation in user_data bits 40-47 left a residual
UAF path. Generations are per-connState-object, so under fd reuse a
closed predecessor's terminal recv/send CQE can collide with the live
occupant's (fd, gen) identity (P = 1/256 per reuse) and be
misattributed at the staleConnCQE chokepoint — decrementing the LIVE
conn's kernelInflight and clearing recvArmed while its own recv is
still kernel-armed. At 1->0 that disarms the conn's close-time cancel
(recvArmed=false skips cancelConnOps, kernelInflight=0 skips
noteClosedInflight) and drainPendingRelease frees the connState with
the kernel still holding a write pointer into cs.buf — exactly the
v1.4.15/7beebb9 corruption class. Reachability is narrow on non-SQPOLL
tiers (io_uring task-work posts the close-path -ECANCELED during the
submit syscall, before the fd can be re-accepted; the window needs the
cancel SQE dropped on a full SQ ring) and wider under SQPOLL, where
completion ordering decouples from the worker's enter.

Bits 48-55 were unused, so widen the tag to bits 40-55:
connState.generation becomes uint16 and genMask / decodeGen /
encodeUserDataGen / encodeConnOpKey follow. 256x fewer collisions
(1/256 -> 1/65536 per fd reuse); previously-encoded values decode
identically since bits 48-55 were always zero.

The misroute is not fully eliminated, so also correct the two comments
that overclaimed the zero clamp (staleConnCQE,
TestLiveTerminalCQEUpdatesInflight): the clamp only protects an
already-drained counter; the residual is now documented in-line, with
the orphaned closedOps entry firing the 5 s backstop WARN as the
production signal when it does fire.

New TestGenNoAliasAcrossOldEightBitBoundary pins that generations
congruent mod 256 (1 vs 257) — aliased pre-widening — encode distinct
user_data and route as stale without touching the live conn's
accounting. Round-trip test extended across the 16-bit range.

Verified: linux/amd64 build+vet+gofmt clean; engine/iouring suite 70/70
PASS on msa2-server (incl. TestAsyncChurnNoUseAfterFree 60 s and
TestAsyncChurnCloseWithStragglerData 30 s).
@FumingPower3925
FumingPower3925 merged commit 584caf2 into main Jun 16, 2026
31 checks passed
@FumingPower3925
FumingPower3925 deleted the milestone/v1.5.0 branch June 21, 2026 14:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment