Skip to content

overload: StageReap calls runtime.GC() on every poll tick, not once on entry #407

Description

@FumingPower3925

Summary

The overload controller's Reap stage forces a full GC on every poll tick while the computed stage is StageReap, not once on the transition into Reap. Under sustained CPU in the Reap band this runs runtime.GC() roughly every PollInterval (default 1s) — a repeated, expensive reclaim precisely when the server is already under load.

Where

middleware/overload/overload.go:259-263, inside the poll loop:

stage.Store(int32(newStage))
if newStage == StageReap && cfg.EnableReap {
    if cfg.ReapAggressiveness >= 1 {
        runtime.GC()
    }
}

newStage == StageReap is true on every tick the signal sits in the Reap band (default CPU 0.80–0.85), so the GC is not one-shot.

Impact

Gated behind EnableReap (opt-in, default off), so only opt-in users are affected — but a forced GC is a one-shot capacity reclaim, and firing it once per second under stable moderate load likely does more harm (GC CPU + STW pauses) than good.

Fix

Fire runtime.GC() only on the entry transition:

if previousStage != StageReap && newStage == StageReap && cfg.EnableReap && cfg.ReapAggressiveness >= 1 {
    runtime.GC()
}

(track previousStage across ticks), or rate-limit forced GCs to a minimum interval.

Related

ReapAggressiveness == 2 ("encourage pool drain", config.go) is documented but behaves identically to 1 — only >= 1 is checked. Worth implementing the level-2 behavior or removing the field.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/engineEngine interface or implementationbugSomething isn't workingperformancePerformance optimization

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions