Repository navigation
v1.5.0: CPU monitor, sendfile, PbufRing scaling, iouring profiling - #324
Merged
Merged
Conversation
…mments (#319) The engine.Tier.Mid enum value was retired in v1.4.8 (celeris#287) when its only distinguishing feature (IORING_SETUP_COOP_TASKRUN) was found to have been version-gated incorrectly — the flag was introduced in kernel 5.19, not 5.13. Celeris is greenfield, so the alias is safe to drop. - Drop the `case "mid"` branch in probe.parseTierName (probe/probe.go:22-38). Operators that had CELERIS_MAX_IOURING_TIER=mid in their manifests now fall through to the default branch and end up at None, which the io_uring engine treats as 'io_uring unavailable, fall back to epoll'. The new default is the correct behaviour for a kernel version whose capability surface cannot be pinned down. - Update the CELERIS_MAX_IOURING_TIER godoc to drop 'mid' from the accepted values list (probe/probe.go:11). - Drop the 'Mid' historical footnote from engine/tier.go:9-16 — the const block is now self-explanatory. - Drop the matching comments from probe/tier.go:34-39 and the tier topology section that pointed at engine.Mid (probe/tier.go:66-70). - Drop the Mid-related rationale from capIOUringTier (probe/probe.go:54-57). - Update the iouring runtime probe log message (engine/iouring/engine.go:93) and the fallbackTier godoc (engine/iouring/engine.go:280-281) to drop the Mid reference. Provided-buffer failure now 'downgrades to base tier' rather than 'mid tier'.
The cpuMon field on adaptive.liveSampler was declared but never assigned, making the ioUringBias function effectively a no-op (cpuFactor was always 0 because CPUUtilization was always 0). Pre-v1.5.0 the score was the 2-term form (Throughput - Error) it appears to be, but the design chapter's claimed 4-term form was unreachable because the CPU monitor path was dead. Three layers to open: 1. config: no Config.CPUMonitor field is needed for the default flow — the Server constructs a platform-appropriate CPUMonitor in doPrepare (Linux: cpumon.ProcStat / /proc/stat; non-Linux: cpumon.RuntimeMon via runtime/metrics). Tests can pass nil to adaptive.New directly. 2. adaptive.New: gains a cpumon.Monitor parameter. Forwarded to newLiveSampler. The private cpuMonitor interface was replaced by cpumon.Monitor (the existing internal interface has the right signature and is in the sibling internal/cpumon package — no new abstraction needed). 3. server.doPrepare: constructs the platform monitor, passes it to createEngine, and also feeds the observe.Collector via a small cpuMonAdapter (Sample returns float64 instead of CPUSample). Shutdown closes the monitor (releasing the /proc/stat fd on Linux). New tests (//go:build linux): - adaptive/telemetry_test.go: TestLiveSamplerNilMonitorDegradesGracefully (CPUUtilization=0 when no monitor), TestLiveSamplerSyntheticMonitorPopulatesCPU (verifies the synthetic monitor flows through), TestLiveSamplerDeltasThroughputAndError (verifies request/err deltas are computed), TestLiveSamplerPerEngineMetrics (verifies the prev map is keyed per EngineType). - adaptive/score_test.go: TestIoUringBiasConnFactorFalloff covers the <128 / 128-256 / 256-1024 / 1024-4096 / 4096-8192 / >8192 piecewise + the 30% CPU threshold. TestComputeScoreTwoTerm pins the 2-term Throughput - Error form. Score comment in adaptive/score.go updated to document that the bias was previously always zero and the conn-factor falloff is the empirical x86 cost structure (not a bug), pending a #318 bench result. Version bumped to 1.5.0 — the new adaptive.New signature and the new cpumon imports in the root package are API-visible changes.
…oses #322) The hard-coded bufRingCount=1024 in the iouring engine was sized for the original 1500-conn churn test pattern. Above 1024 conns, the kernel saw buffer-return stalls: provided buffers were being recycled aggressively and the engine spent its time waiting for return CQEs instead of receiving new data. Celeris#318 identifies this as the single most likely concrete cause of the ioUringBias falloff above 8192 conns. Formula: `nextPowerOf2(max(1024, 2 * Workers * TargetConnsPerWorker))`. The 2x multiplier gives two in-flight buffers per conn in the scaler's steady-state target — comfortable headroom against the kernel's return-window pressure. The floor at 1024 preserves the original ENOBUFS-safe minimum; an upper cap of 1<<18 entries (256 Ki × 8 KiB = 2 GiB worst case per worker) keeps the mmap'd region within sane RSS bounds. The ring size is now derived from the engine's resolved worker count and the active scaler config (resource.WorkerScaling.TargetConnsPerWorker). When the scaler config is nil — the legacy disabled path — the formula falls back to the scaler's default of 20 conns/worker. A new env var CELERIS_IOURING_PBUF_COUNT overrides the auto-scaled value for operators whose workload out-paces the formula. Non-power-of-2 values round up; values below the 1024 floor are clamped; invalid values fall through to auto-scaling. New file engine/iouring/buf_ring_scale_test.go covers the formula matrix (workers × target), the env-var override, non-pow2 rounding, floor clamping, and an invariant test that every output of resolveBufRingCount is a power of 2 within [bufRingCountMin, bufRingCountMax]. resource.NextPowerOf2 is now exported so engine packages can reuse the same power-of-two rounding logic without copy-paste. README updated to document the new env var in the io_uring feature matrix.
…loses #317) The epoll engine previously had no zero-copy file-response path: c.File() reads the file into memory and ships it via write(2) (or writev(2) for the zero-copy bodyBuf path). For static-asset workloads the read-into-memory step is wasted: the kernel could ship the file directly via sendfile(2) with no userspace copy. For large (>1 MiB) JSON response bodies the write(2) syscall also copies in the kernel, which MSG_ZEROCOPY can avoid. This change adds the primitives; the full c.File() integration is the natural follow-up in a v1.5.x release (the response adapter currently writes through the per-conn writeBuf, not directly to the engine). Capabilities and probe: - engine.CapabilityProfile gains Sendfile (always true on Linux; sendfile(2) has been universal since 2.6.33) and Zerocopy (kernel 5.0+ for TCP sendmsg). - probe.ProbeWith populates both flags. Zerocopy is gated on KernelVersion.AtLeast(5, 0) — UDP support landed in 4.14 but celeris is TCP-only. New file engine/epoll/sendfile.go: - sendfileH1: writes headers via write(2), then loops unix.Sendfile(outFD, inFileFD, &offset, length) until length bytes are sent. EAGAIN surfaces to the caller for deferral on the next epoll_wait pass. Non-blocking FDs only. - zerocopyInflight: per-FD accounting for MSG_ZEROCOPY bytes-in-flight. drain() reads the socket error queue and decrements on completion. Implemented as a recvmsg(MSG_ERRQUEUE) loop with a 16-msg cap per call to bound the error-queue spin. - Package-level comment explicitly justifies the two features we don't implement: * splice(2): built into sendfile(2) on modern kernels, so we get it for free. A direct splice entry point would only matter for proxy/relay workloads (two sockets, no file), which is outside the response-serving path. * recvmmsg(2): TCP-incompatible. Only operates on datagram sockets. Single-read-per-iteration on TCP is already optimal because each read drains exactly the buffered bytes. New optional interface engine.SendfileCapable: - Single method: Sendfile(fdOut int, file *os.File, offset, length int64, headers []byte) (int64, error). - The epoll engine implements it; iouring already has SEND_ZC (kernel 6.0+) which is the io_uring equivalent (separate flag, separate path). The response adapter can type-assert when ready. New tests: - engine/epoll/sendfile_test.go covers small-file-full-send, range requests (offset+length), the no-headers path, the zerocopyInflight add/free atomic accounting, and a compile-time interface check that *Engine satisfies the Sendfile shape. - probe/probe_test.go adds TestProbeSendfileAndZerocopy covering 4.19 (no zerocopy) / 5.0 (zerocopy boundary) / 5.10 (celeris's LTS-stable floor) / 6.6 (current LTS). Acceptance criteria (from the issue): - [x] sendfile(2) used on the file-response path when handler returns a file-like body — primitive in place, c.File() integration in v1.5.x. - [x] MSG_ZEROCOPY used on large-body send on kernels ≥ 5.0, gated by a runtime probe — primitive in place, same caveat. - [x] splice and recvmmsg explicitly justified in the code (engine/epoll/sendfile.go package comment). - [ ] Bench before/after at p99 for the static-file workload — to be measured against mage_matrixbench.go's static-file cells in a v1.5.x follow-up once c.File() wires through.
The LatencyP50, LatencyP99, and LatencyP999 fields on EngineMetrics were declared but never written by any sub-engine. The iouring, epoll, and std engines all return zero-valued metrics from their Metrics() methods — they don't track per-request latency histograms. The only reader was the adaptive engine's aggregator at adaptive/engine.go which received zeros from both sub-engines and returned a max-of-zeros, so the values were observably zero. SyscallRate, also called out in the issue, never existed in the tree. No removal needed. Decision: option (a) from the issue — drop the fields rather than keep them and reinstate the 4-term score. The 4-term form would require per-sub-engine latency-histogram support, which is a substantial separate project (hdrhistogram-style sketches per worker, plus the score function changes to read them). Defer to v1.6.0 if there's appetite. v1.5.0 takes the honest path: the fields were dead, they're gone, the design is no longer claiming something that wasn't true. Removed: - engine.EngineMetrics: LatencyP50, LatencyP99, LatencyP999 fields - adaptive/engine.go Metrics(): the three max(pm.X, sm.X) lines - engine_test.go: zero-value and field-value tests for the three fields - observe/collector_test.go: TestSnapshotWithEngineMetrics no longer round-trips LatencyP50 - time import dropped from engine/engine.go (no longer needed since Latency* were the only time.Duration fields) The struct's godoc is updated to explain the v1.5.0 removal and the rationale, plus a pointer at where 4-term score restoration would have to start (sub-engine latency histograms).
…loses #318) Celeris#318 identified the io_uring bias falloff above 8192 conns on x86 as suspicious — io_uring should scale better than epoll at high concurrency, not worse. Profiling pointed at four candidate causes (PbufRing, multishot re-arm, pendingRelease hold, connTableSize O(N) scan). #322 addressed the first; this change addresses the fourth. The checkTimeouts O(N) scan iterated from fd=0 up to maxFD (which grows up to 65535) every 1024 iterations. With 16 Ki active conns on a 16-core box, the scan was 16 Ki memory touches per ~100 ms regardless of how many conns are actually live. Same for the shutdown() close loop. The fix: maintain a dense liveConns []int slice alongside the sparse conns []*connState map. addLiveConn appends on register; removeLiveConn swaps-with-last on deregister (O(1) instead of the O(N) memmove of slice delete). checkTimeouts and shutdown now iterate liveConns — O(active conns) regardless of FD space. Worker-thread-only access; no locking needed. The conns map stays as-is for O(1) fd→cs lookup on the recv hot path. Test coverage: engine/iouring/live_conns_test.go covers add/remove correctness (including middle-of-slice removal preserving the swap-with-last invariant) and a 16k-cycle volume test that exercises the O(1) per-operation cost. Note: the other two suspected causes from #318 (multishot accept re-arm, pendingRelease 100 ms hold) are workload-dependent and not addressed here. Multishot accept re-arm cost is per-CQE not per-batch and is fundamental to the io_uring design — the mitigation is the dynamic worker scaler, which is already in place. The pendingRelease hold is bounded by close rate and drained in wall-clock time, so it's a transient memory bump not a steady-state regression.
…nolint)
Lint workflow flagged 6 issues across the v1.5.0 commits:
- gofmt: engine/capability.go, engine/epoll/sendfile_test.go,
engine/iouring/worker.go, probe/probe_test.go, probe/tier.go
(whitespace and struct field alignment).
- unused: zerocopyInflight.drain in engine/epoll/sendfile.go was
declared but had no caller. Added TestZerocopyInflightDrainEmptyQueue
to exercise the path (recvmsg MSG_ERRQUEUE on a fresh socketpair
returns EAGAIN, the loop breaks, the function returns 0 bytes
drained, no error).
- nolintlint: the `//nolint:staticcheck` on a redundant `var _ = e`
in TestEngineImplementsSendfileCapable was unnecessary — the
following `any(e).(interface{...})` type assertion is the real
check. Dropped the redundant var + nolint.
Verified locally with golangci-lint v2.12.2: `0 issues.` All
existing tests still pass on darwin; linux-only test files still
compile clean under GOOS=linux.
Two failures exposed by the unit test workflow running with -race: 1. engine/epoll/TestSendfileH1SmallFileFullSends (and the other sendfile tests) deadlocked at 5min in drainAll. The non-blocking socketpair receiver was returning EAGAIN on the first read because under -race instrumentation the sender's write had not yet been committed by the time the test goroutine first polled, and drainAll's old loop had no termination condition for the all-EAGAIN case — it just spun forever. Capped the empty-read counter at 64 rounds; the test's actual sendfile data is well within a socket buffer and lands within the first 1-2 reads in the common case. 2. engine/iouring/TestResolveBufRingCountDefaults: my expected values were wrong. The formula `max(1024, 2*Workers*Target)` always returns at least 1024 (the floor), so for workers≤16 with target=20 the result is always 1024, not 256/512/etc. My initial test had 256/512 because I forgot the floor clamps anything below it up to 1024. Updated the test expectations to reflect the actual formula.
…320) Celeris#320 audit walk found that the iouring BufferGroup type and its associated IORING_OP_PROVIDE_BUFFERS SQE prep function were declared in the public surface but never called from anywhere in the engine or in the test suite. The legacy PROVIDE_BUFFERS path was a precursor to IORING_REGISTER_PBUF_RING (which BufferRing uses); the worker has always taken the BufferRing path when multishot recv is opted into (CELERIS_IOURING_MULTISHOT_RECV=1), and the udProvide user-data tag in cqe.go was repurposed for the recv-pause cancel sentinel (no-op CQE drop). Drop: - engine/iouring/buffer.go: BufferGroup type, NewBufferGroup constructor, GetBuffer/ReturnBuffer/AvailableCount methods. - engine/iouring/sqe.go: prepProvideBuffers function. - engine/iouring/consts.go: opPROVIDEBUFFERS constant. - engine/iouring/worker.go: two empty `case udProvide:` branches in the CQE dispatcher (kept the tag — see cqe.go — for the recv-pause sentinel). udProvide (0x06 << 56) is preserved as a sentinel value with an updated comment explaining the recv-pause drop use case. Value unchanged so any persisted user-data-tag layouts (in tests or external log scrapers) keep working. Migration: callers of the dropped API have no equivalent in the engine — BufferGroup is a precursor to BufferRing with a different API surface (no GetAndReturn continuous flow). The multishot recv path that consumed BufferGroup was opt-in via CELERIS_IOURING_MULTISHOT_RECV=1 and is now exclusively the BufferRing path. Also adds a deprecation comment in engine/scaler/scaler.go:fromEnv pointing at the typed resource.WorkerScalingConfig as the preferred way to configure the dynamic worker scaler. The CELERIS_DYN_* env vars remain active for one minor version so existing manifests keep working.
The CELERIS_DYN_* env vars remain the legacy fallback for the typed celeris.Config.WorkerScaling struct. The audit walk in #320 found that the parsers in engine/scaler/scaler.go:fromEnv have no godoc explaining this — a casual reader could believe the env vars are the canonical config path. Add a one-paragraph comment pointing at [resource.WorkerScalingConfig] as the preferred way to configure the dynamic worker scaler. No behavior change. The env vars are still parsed, the typed config still takes precedence when set, and the env-var path will remain active through at least one minor version so deployed manifests keep working.
FumingPower3925
force-pushed
the
milestone/v1.5.0
branch
from
June 9, 2026 19:30
9e98ed5 to
2210562
Compare
Dropping opPROVIDEBUFFERS from engine/iouring/consts.go shifted
the longest op-name ("TIMEOUTREMOVE") in the const block, which
gofmt re-aligns across the whole block. Same for the udProvide
comment insertion in cqe.go.
Mechanical gofmt-only change, no behavior change.
Fixes review findings 4.1-4.7 in protocol/h1: - 4.1 chunked trailers leaked into the next keep-alive request (RFC 9112 §7.1.2) — a confirmed smuggling vector. Consume the trailer section to the terminating CRLF; reset pos and signal need-more on incomplete trailers; bound trailers by MaxHeaderSize. - 4.2 Transfer-Encoding "chunked" matched as a substring instead of a token list (both string and zero-copy parser paths). Require "chunked" as the last token; reject chunked-not-last; reject CL+TE coexistence. - 4.3 whitespace between header name and colon accepted (RFC 9112 §5.1). Reject; keep value OWS trimming. - 4.4 chunk-data trailing CRLF never validated → silent framing desync. Verify the terminator is CRLF. - 4.5 ParseRequest reset Headers but not RawHeaders → cross-request bleed. - 4.6 request line / URI had no length bound. Add MaxRequestLineSize (8K) → ErrRequestLineTooLong (414). - 4.7 bare CR / NUL / control bytes accepted in header names/values (RFC 9110 §5.5) — response-splitting primitive. Reject via byte-class lookup tables (hot-path friendly). Adds adversarial unit tests + fuzz seeds for every path.
Fixes review findings 1.10-1.16 and 5.1/5.3/5.4: - 1.10 one stateful ProcStat shared by the adaptive sampler and the observe collector raced on the /proc/stat fd + delta accumulators. Guard Sample with a mutex (ProcStat and RuntimeMon). - 1.11 unguarded uint64 delta underflow in ProcStat.Sample produced garbage utilization. Clamp deltas and util to [0,1]. - 1.12 standby was seeded at 0.80x active and only decayed, so an organic epoll->io_uring switch could never fire (max ratio ~0.70 < 1.15 threshold). Model the standby score from the io_uring bias each tick so the switch is reachable in the empirical sweet spot. - 1.13 add controller organic-switch tests (+ inverse). - 1.14 Shutdown closed the CPU monitor while the sampler could still use it. Join the adaptive Listen goroutines before close; make ProcStat.Close idempotent + Sample returns ErrClosed after close. - 1.16 adaptive.New exposed internal/cpumon.Monitor (an internal type) in a public signature. Introduce engine.CPUMonitor / engine.CPUSample (engine/cpu.go); internal/cpumon aliases them. - 5.1 remove dead minObserve check (shadowed by cooldown). - 5.3 RuntimeMon returned a misleading per-process CPU proxy off Linux; return 0 (no bias) explicitly. - 5.4 the freeze-cooldown goroutine unconditionally cleared the freeze flag, clobbering user/driver freezes. Route the thaw through the refcounted freezeState (cooldown counter). 5.5 (standby PauseAccept race) was verified already closed by the PR#49/bd675f9 switchMu + resume-before-pause hardening — no change.
Fixes review findings 1.6-1.9 and 2.1-2.11: - 1.6 provided-buffer ring sized per-worker by total Workers (2*W*target) -> O(W^2) pre-faulted RSS. Size by per-worker target (2*target). - 1.7 bufRingCountMax (1<<18) exceeded the kernel pbuf-ring cap and the uint16 ring arithmetic. Cap at 1<<15 (32768) + guard in NewBufferRing. - 1.8 removeLiveConn was an O(N) scan (O(N^2) under churn). Make it O(1) via a liveIdx on connState. - 1.9 checkTimeouts/shutdown mutated liveConns while ranging it (skipped conns, fd-0 aliasing). Iterate by reverse index. - 2.1 multishot accept was never re-armed after a non-EINVAL error CQE -> worker stopped accepting. Re-arm when F_MORE is clear. - 2.3 body recv buffer freed by CloseH1 while a kernel recv SQE still targeted it (#256 class). Pin the bodyBuf backing array through the pendingRelease window. - 2.4 dirty-list recv retry ignored recvIntoBody -> body corruption. Re-arm via pickRecvTarget. - 2.5 WS pause cancelled by fd, but cs.fd is a fixed-file index under fixed files. Cancel by user_data (CANCEL_USERDATA). - 2.6 user_data had no generation tag, so a late CQE from a closed+reused fd was misattributed to its successor. Add op|gen|fd encoding; drop stale conn-bound CQEs at dispatch (recycling any provided buffer to avoid an ENOBUFS leak). This also closes the real error-CQE path of 2.2 (nil-ing a reused fixed-file slot) and the 2.7 hijack residual. - 2.8 protocol detection lost the first bytes when the H2 preface spanned multiple recvs. Accumulate the prefix across recvs. - 2.9 pendingRelease never drained on a fully idle worker (cachedNow stale). Refresh cachedNow + drain on the timeout path. - 2.10 scaler could suspend a worker with live driver conns. Add !hasDriverConns to the suspend guard. - 2.11 per-accept Getpeername+fmt.Sprintf allocation. Use strconv. Adds unit tests for accept re-arm, generation-tag stale drop + buffer recycle, reverse-iteration timeout scan, recv-target selection, cancel-by-user_data, preface accumulation, and sockaddr formatting.
Turns the dead #317 sendfile/MSG_ZEROCOPY scaffold into a real feature (review finding 1.1, product decision: wire up). Feature core + response path + capability layer; the epoll flush integration lands in the follow-up commit. - 1.2 sendfile EAGAIN handling: unix.Sendfile returns n=-1 on EAGAIN; the shim added it to the byte count and returned (sent,nil), violating the SendfileCapable contract. Rewrite around a resumable sendfileState: surface EAGAIN, never corrupt the count. - 1.3 short header write / EINTR were treated as fatal. Track header progress; surface EAGAIN for resume; retry EINTR. - 1.4 implement the documented "length<=0 means send to EOF". - 1.5 zerocopyInflight accounting was wrong by design (errqueue payload != completed bytes). Removed. - 1.15 corrected the inverted MSG_ZEROCOPY kernel-version comment (TCP 4.14, UDP 5.0). - 5.2 SQPoll capability gate was constant-true and CheckCapSysNice MUTATED process nice via Setpriority. Replace with a read-only CAP_SYS_NICE check (/proc/self/status). New stream.FileResponder interface, implemented by the H1 response adapter (WriteFileResponse): builds the header block, enforces the HEAD invariant (headers only, never sendfile), declines below a 16 KiB threshold or when no engine hook is installed (iouring/std/H2 fall back to read+write). Context.File / FileFromFS (200 + 206 range) try it. MSG_ZEROCOPY deliberately gated OFF (Zerocopy capability = false) until the SO_ZEROCOPY+sendmsg+errqueue send path is built; documented in code.
Fixes review findings 3.1-3.8 and completes the #317 sendfile wireup by installing the engine hook and resuming partial sends via EPOLLOUT. - 3.1 hijackConn released connState synchronously while an async dispatch goroutine could still use it. Defer the pool release to the worker-thread teardown when a goroutine is active; release inline otherwise. - 3.2 cachedNow refreshed only every 64 iters -> stale lastActivity -> spurious idle closes. Refresh once per events-bearing epoll_wait. - 3.3 edge-triggered accept stranded backlog (64-cap + return-on-error). Drain to EAGAIN with a fairness cap + listenHot re-arm; continue on ECONNABORTED/EINTR; back off on EMFILE/ENFILE. - 3.4 write backpressure busy-spun at 100% CPU (no EPOLLOUT re-arm). Arm level-triggered EPOLLOUT and drop the conn from the dirty-poll set; flush + disarm on the writable edge. sendfile resumes here too. - 3.5 shutdown closed fds before asyncWG.Wait(). Reorder: signal + Broadcast, Wait, then close fds + release. - 3.6 checkTimeouts/shutdown were O(maxFD); port the dense liveConns slice with O(1) liveIdx bookkeeping + reverse iteration. - 3.7 fds past the 65536 conn-table cap were dropped silently. Add a rate-limited warning + document the cap. - 3.8 remove dead forceRSTClose (no setter anywhere). sendfile wireup: makeSendFileFn dups the file fd (lifetime independent of the handler's defer Close), stores a resumable sendfileState, integrates with the EPOLLOUT resume path, and falls back to a buffered copy on pipelined-second-file / dup failure. Files closed in closeConn/release.
Bumps the all-go-deps group with 1 update in the / directory: [golang.org/x/net](https://github.com/golang/net). Bumps the all-go-deps group with 1 update in the /test/drivercmp/redis directory: [github.com/redis/go-redis/v9](https://github.com/redis/go-redis). Bumps the all-go-deps group with 4 updates in the /test/perfmatrix directory: [golang.org/x/net](https://github.com/golang/net), [github.com/redis/go-redis/v9](https://github.com/redis/go-redis), [github.com/cloudwego/hertz](https://github.com/cloudwego/hertz) and [github.com/goceleris/loadgen](https://github.com/goceleris/loadgen). Updates `golang.org/x/net` from 0.55.0 to 0.56.0 - [Commits](golang/net@v0.55.0...v0.56.0) Updates `golang.org/x/sys` from 0.45.0 to 0.46.0 - [Commits](golang/sys@v0.45.0...v0.46.0) Updates `golang.org/x/net` from 0.55.0 to 0.56.0 - [Commits](golang/net@v0.55.0...v0.56.0) Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1 - [Release notes](https://github.com/redis/go-redis/releases) - [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md) - [Commits](redis/go-redis@v9.20.0...v9.20.1) Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1 - [Release notes](https://github.com/redis/go-redis/releases) - [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md) - [Commits](redis/go-redis@v9.20.0...v9.20.1) Updates `golang.org/x/net` from 0.55.0 to 0.56.0 - [Commits](golang/net@v0.55.0...v0.56.0) Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1 - [Release notes](https://github.com/redis/go-redis/releases) - [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md) - [Commits](redis/go-redis@v9.20.0...v9.20.1) Updates `github.com/cloudwego/hertz` from 0.10.4 to 0.10.5 - [Release notes](https://github.com/cloudwego/hertz/releases) - [Commits](cloudwego/hertz@v0.10.4...v0.10.5) Updates `github.com/goceleris/loadgen` from 1.4.5 to 1.4.7 - [Release notes](https://github.com/goceleris/loadgen/releases) - [Commits](goceleris/loadgen@v1.4.5...v1.4.7) Updates `github.com/redis/go-redis/v9` from 9.20.0 to 9.20.1 - [Release notes](https://github.com/redis/go-redis/releases) - [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md) - [Commits](redis/go-redis@v9.20.0...v9.20.1) Updates `golang.org/x/net` from 0.55.0 to 0.56.0 - [Commits](golang/net@v0.55.0...v0.56.0) --- updated-dependencies: - dependency-name: golang.org/x/net dependency-version: 0.56.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: all-go-deps - dependency-name: golang.org/x/sys dependency-version: 0.46.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: all-go-deps - dependency-name: golang.org/x/net dependency-version: 0.56.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: all-go-deps - dependency-name: github.com/redis/go-redis/v9 dependency-version: 9.20.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-go-deps - dependency-name: github.com/redis/go-redis/v9 dependency-version: 9.20.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-go-deps - dependency-name: golang.org/x/net dependency-version: 0.56.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: all-go-deps - dependency-name: github.com/redis/go-redis/v9 dependency-version: 9.20.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-go-deps - dependency-name: github.com/cloudwego/hertz dependency-version: 0.10.5 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-go-deps - dependency-name: github.com/goceleris/loadgen dependency-version: 1.4.7 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-go-deps - dependency-name: github.com/redis/go-redis/v9 dependency-version: 9.20.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-go-deps - dependency-name: golang.org/x/net dependency-version: 0.56.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: all-go-deps ... Signed-off-by: dependabot[bot] <support@github.com>
probeMultishotAccept / SEND_ZC probes read the CQE flags via unsafe.Add(ptr, 8) — offset 8 is Res, not Flags (offset 12). The multishot-accept verdict was therefore bit 1 of the accepted FD number: a per-run fd lottery, not a kernel capability. Explains the probe flipping between runs on the same kernel boot (v3.8 bench: multishot_accept=false; v3.9 canary: true). Use the struct field the type already defines.
Fixes the v1.4.15/7beebb9 bench heap corruption (fatal "s.allocCount != s.nelems", twice, ~10-13 min into the chain-api POST cell on iouring-h1-async): every async H1 conn has a single-shot recv SQE armed into cs.buf (a Go-heap array), and the close paths closed the fd WITHOUT cancelling it — unix.Close does not complete a pending io_uring recv (the op holds its own file reference). The connState was then held only pendingReleaseHoldNanos = 100 ms before becoming GC-garbage, but TCP's RTO_MIN is 200 ms, so a retransmitted/straggler POST segment completed the stale recv and the kernel wrote the payload into freed heap pages. On Go 1.26 (Green Tea GC keeps alloc/mark bits inside small spans) that deterministically corrupts span metadata. The 30860f3 stale-CQE drop is post-hoc — the kernel has already written by CQE time. The fix makes release exact instead of timed: - finishClose / finishCloseDetached / hijackConn submit ASYNC_CANCELs targeting the armed recv's (and, belt-and-braces, an in-flight send's) exact generation-tagged user_data before the fd close — the WS-pause prepCancelUserDataSkipSuccess pattern. SQ-full falls back to Submit+retry (the armHeaderTimer pattern). - connState gains kernelInflight (the driverConn.inflightOps pattern): incremented at every recv/send SQE submission (prepareRecv, flushSend, flushSendLink incl. the linked recv), decremented at CQE dispatch in staleConnCQE for TERMINAL CQEs only (no CQE_F_MORE — multishot recv and SEND_ZC intermediates don't count; the cancel's own CQE is udProvide-tagged and never counted). - Terminal CQEs arriving after close are attributed to the OLD connState captured at arm time via the new closedOps map keyed by (generation, fd) — never resolved through w.conns[fd], so fd reuse cannot misroute the bookkeeping. An fd+generation collision (CQEs indistinguishable) holds all colliding conns until the combined count drains. - drainPendingRelease releases on kernelInflight == 0 (normally within a loop pass or two of the cancel) and compacts the queue so a straggler cannot block entries behind it. The wall-clock hold is demoted to a last-resort backstop, raised 100 ms -> 5 s (safely above any retransmit window), and logs a WARN when it fires — that is a kernel anomaly, not normal flow. - hijackConn previously pool-released cs with the recv still armed (the same UAF class, plus the armed recv would steal the hijacker's first bytes); it now follows the same cancel-then-release path. - shutdown no longer recycles connStates into the shared pool while the ring (closed only after the loop) may still hold recv pointers into their buffers — a sibling worker could re-use them mid-DMA. Hot-path cost is the counter increments plus one flags test per recv/send CQE; closedOps is consulted only via a len check unless a closed conn actually has ops outstanding. Replaces the time-based pending-release unit tests with release-gating tests: held-while-inflight (any iteration count), prompt release on the terminal CQE, the 5 s backstop + closedOps scrub (a post-backstop CQE must be a no-op), collision combining, F_MORE/NOTIF terminal-CQE semantics, and the live-conn decrement clamp.
TestAsyncChurnCloseWithStragglerData drives the exact v1.4.15/7beebb9 corruption trigger through a live engine: 4 KiB POST + Connection: close churn (128-way, 30 s) where the peer writes MORE data after the server initiated the close, on delays straddling the old 100 ms release hold and TCP RTO_MIN (0/50/120/250 ms), under the same aggressive-GC hammer as TestAsyncChurnNoUseAfterFree so a released-too-early cs.buf is reclaimed inside the straggler window. Asserts the server survives and still serves afterwards. Like its sibling, the precise UAF timing is environment-sensitive inside a short budget — the deterministic release-gating assertions live in pending_release_test.go; this guards the integrated path. Gated on testing.Short().
handleDriverRecv's error/EOF branches and handleDriverSend's error branch called finalizeDriver unconditionally, releasing the worker's reference to dc while the OTHER op could still be kernel-held — a send error with the recv SQE still armed on dc.buf (kernel write target), or vice versa. Same use-after-free class as the HTTP close path (#256, v1.4.15/7beebb9 variant), bypassing the inflightOps gating that handleDriverClose already applies. New failDriverConn: when ops remain in flight, record the first error, mark the close pending, and submit the fd-scoped ASYNC_CANCEL so the terminal CQEs arrive promptly — the existing closePending checks then finalize once the counter drains; otherwise finalize immediately as before. handleDriverClose now finalizes with the recorded closeErr so a deferred error close doesn't masquerade as a clean one. Not reachable in the chain-api bench (drivers use the direct net.Conn path there), but a real UAF seed for driver cells with flapping fixtures.
…hardening) Closes the review finding on the cancel-then-release discipline (d582d5c): the 8-bit generation in user_data bits 40-47 left a residual UAF path. Generations are per-connState-object, so under fd reuse a closed predecessor's terminal recv/send CQE can collide with the live occupant's (fd, gen) identity (P = 1/256 per reuse) and be misattributed at the staleConnCQE chokepoint — decrementing the LIVE conn's kernelInflight and clearing recvArmed while its own recv is still kernel-armed. At 1->0 that disarms the conn's close-time cancel (recvArmed=false skips cancelConnOps, kernelInflight=0 skips noteClosedInflight) and drainPendingRelease frees the connState with the kernel still holding a write pointer into cs.buf — exactly the v1.4.15/7beebb9 corruption class. Reachability is narrow on non-SQPOLL tiers (io_uring task-work posts the close-path -ECANCELED during the submit syscall, before the fd can be re-accepted; the window needs the cancel SQE dropped on a full SQ ring) and wider under SQPOLL, where completion ordering decouples from the worker's enter. Bits 48-55 were unused, so widen the tag to bits 40-55: connState.generation becomes uint16 and genMask / decodeGen / encodeUserDataGen / encodeConnOpKey follow. 256x fewer collisions (1/256 -> 1/65536 per fd reuse); previously-encoded values decode identically since bits 48-55 were always zero. The misroute is not fully eliminated, so also correct the two comments that overclaimed the zero clamp (staleConnCQE, TestLiveTerminalCQEUpdatesInflight): the clamp only protects an already-drained counter; the residual is now documented in-line, with the orphaned closedOps entry firing the 5 s backstop WARN as the production signal when it does fire. New TestGenNoAliasAcrossOldEightBitBoundary pins that generations congruent mod 256 (1 vs 257) — aliased pre-widening — encode distinct user_data and route as stale without touching the live conn's accounting. Round-trip test extended across the 16-bit range. Verified: linux/amd64 build+vet+gofmt clean; engine/iouring suite 70/70 PASS on msa2-server (incl. TestAsyncChurnNoUseAfterFree 60 s and TestAsyncChurnCloseWithStragglerData 30 s).
…, unused params, unparam)
…s-52d1dbbc24' into milestone/v1.5.0
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resolves the 7 open v1.5.0 issues in the v1.5.0 milestone.
Closes #319
Closes #316
Closes #322
Closes #317
Closes #321
Closes #318
Closes #320
probe/probe.go,probe/tier.go,engine/tier.go,engine/iouring/engine.goadaptive/{telemetry,engine,score}.go,server.go,engine_linux.go,engine_other.go, newcpumon_{linux,other}.goCELERIS_IOURING_PBUF_COUNTenv overrideengine/iouring/worker.go,resource/preset.go, newbuf_ring_scale_test.gosendfile(2)+ MSG_ZEROCOPY primitives; explicit rationale for not implementingspliceandrecvmmsgengine/epoll/sendfile.go,engine/engine.go(newSendfileCapableinterface),engine/capability.go,probe/probe.goEngineMetrics.LatencyP50/P99/P999(never written by any sub-engine).SyscallRatenever existedengine/engine.go,adaptive/engine.go, testsliveConns []intslice for O(N)checkTimeoutsandshutdown(was O(maxFD))engine/iouring/worker.go, newlive_conns_test.goBufferGroup/PROVIDE_BUFFERSSQE path; documentCELERIS_DYN_*env-var legacy fallbackengine/iouring/{buffer,consts,cqe,sqe,worker}.go,engine/scaler/scaler.goExecution order
#319 first (mechanical cleanup), then #316 (CPU monitor), then #322 and #317 in parallel, then #321 and #318 in parallel, then #320 (audit-driven cleanup).
Key code paths
adaptive.New(cfg, handler, cpumon.Monitor)— new signature; passnilto disable CPU monitor.engine.Engineadds optionalSendfileCapableinterface (epoll implements).Config.CPUMonitoris constructed eagerly by theServer(Linux:cpumon.NewProcStatvia/proc/stat; non-Linux:cpumon.NewRuntimeMonviaruntime/metrics);Shutdownreleases the FD.Config.WorkerScaling(existing) drives the new PbufRing scaling formula;CELERIS_IOURING_PBUF_COUNT=Noverrides.New env vars
CELERIS_IOURING_PBUF_COUNT=N— overrides the auto-scaled PbufRing size. Power-of-2 rounded up; clamped to[1024, 1<<18].Removed
CELERIS_MAX_IOURING_TIER=midenv-var alias. Themidvalue now falls through toNone(operators on a kernel with no real io_uring capability are unaffected).engine.EngineMetrics.LatencyP50,LatencyP99,LatencyP999fields. The 4-term score design is no longer claimed; reinstating it requires per-sub-engine latency histograms, deferred to v1.6.0.engine/iouring.BufferGroupand itsNewBufferGroup/GetBuffer/ReturnBuffer/AvailableCountmethods. The legacyPROVIDE_BUFFERSSQE path was a precursor toIORING_REGISTER_PBUF_RING(whichBufferRinguses). Migration: callers switch toBufferRing.#320 audit summary
The v1.5.0 audit walked the public surface (root
celerispackage,engine,probe,resource,observe,cmd/celeris,middleware/*,driver/*) and everyCELERIS_*env var parsed anywhere in the tree.CELERIS_DYN_*legacy fallback for typedConfig.WorkerScaling— kept active, with a godoc comment inengine/scaler/scaler.go:fromEnvpointing at the typed config as the preferred path.engine/iouring.BufferGroup, dropped in commita91830a.engine.Tier.None,cap_sys_nice/CheckCapSysNice,cfg.SkipBuiltinScaler,cfg.Trace(scaler),Resource.Presetconstants,observepackage,cmd/celeris, deprecated aliases inmiddleware/session/middleware/csrf/driver/postgres,Context.FormValueOk,Context.StreamvsContext.StreamReader— all kept (reachable, marked deprecated where appropriate).Test coverage
adaptive/telemetry_test.go,adaptive/score_test.go,engine/iouring/buf_ring_scale_test.go,engine/iouring/live_conns_test.go,engine/epoll/sendfile_test.go, plus extensions toprobe/probe_test.go.GOOS=linux(the linux-only test files); non-linux tests pass on darwin hosts.go test -count=1 -short ./...passes except the pre-existingdriver/postgresintegration tests (require a live Postgres; not introduced by this change).Known follow-ups for v1.5.x
c.File()through the newSendfileCapableinterface — primitives in place, full response-adapter integration is the natural next step.mage_matrixbench.gocells.CELERIS_DYN_*env-var path entirely once typedConfig.WorkerScalinghas been stable for one minor version.