Skip to content

Keep recurring collectors on a fixed grid instead of drifting from the sweep body start - #4644

Merged
erikdarlingdata merged 2 commits into
devfrom
fix/4636-collector-grid-cadence
Sep 28, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
fix/4636-collector-grid-cadence

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Fixes #4636, Fixes #4640

What changed

The sweep set each due collector's next due time to now + interval, where now is when the server's sweep body started. Every late start pushed the schedule back, so busy stores collected about every 71 seconds instead of 60. A new pure helper in the shared collectors project, CollectorCadence.NextDue(due, now, interval) (used by Darling and Lite), returns the previous due time plus the interval. If that slot is already past, it skips the missed slots and returns the next grid slot, so a stalled server resumes on its grid with no catch-up burst. RunDueCollectorsAsync uses it. Comments on the advance and on ComputeSeededNextDue now say the steady-state advance is grid-based.

Interval changes: the schedule reload re-seeds or caps NextDue when the effective interval changes (DarlingWorker.cs ~4692–4709), and the helper reads the interval per call, so a change takes effect on the next advance.

Lite (#4640): the loop waited a full minute after each cycle's work, and the due check compared the current time with the collector's own start time, so a 1-minute collector missed cycles whenever a cycle ran long. CollectionBackgroundService now keeps a logical cycle time on a fixed grid (CollectorCadence.NextDue, waiting only the remainder to the next slot). RunDueCollectorsAsync(cycleStartUtc, ...) passes that time to a new ScheduleManager.GetDueCollectorsForServer(serverId, atUtc) overload and records each run at the cycle time (RunCollectorAsync(server, name, scheduledAtUtc, ct) -> MarkCollectorRunForServer). The old overloads keep their meaning (they use the current time). The due check now compares exact grid times: a 1-minute collector is due every cycle, a 5-minute one every fifth.

Field read on one production store before the change: 1-minute collectors had p50 71.5-71.9 s between runs (~50 runs/server/hour); 5-minute collectors p50 ~307 s (10.4 runs/hour).

The volume effect. Restoring the configured cadence means more collection. On that store, 1-minute collectors go from ~50 to 60 runs per server per hour (+20%) and 5-minute collectors from ~10.4 to 12 (+15%), so the per-minute tables take ~15–20% more rows and store bytes per day, and the store's WAL rises with them. The sweep gate's busy share rises in step: on that store from 49% to ~59% at width 6.

A manual "collect now" run in Lite (RunCollectorAsync without a cycle time) still records its own start time. After one, that collector's next due check compares the next cycle time with a wall-clock time, so it can skip one cycle. That's the same as today, and only after a manual run.

Pins (DarlingSweepSchedulingTests, Lite.Tests/CollectorGridScheduleTests)

  • on-time run: due + 60 s; run 14 s late: still due + 60 s; after a 5-minute stall: next grid slot, no burst; exact boundary: following slot.
  • Tick simulation: 15 s ticks, body start 0–12 s late (seeded), 1-minute collector, 60 simulated minutes. Grid: 60 runs. The old now + interval rule: 53 runs. The pin asserts 59–61 and that the old rule is lower.
  • 5-minute tick simulation over 60 minutes: grid gives exactly 12 runs; the old rule never more. (At this jitter the old rule loses nothing for a 5-minute collector, since a slot is 20 ticks wide; the field read shows the drift is real on busy stores.)
  • Lite-shaped loop (cycle work 0-20 s, seeded): new shape 60 one-minute runs and 12 five-minute runs per hour; the old shape (wait a minute after the work, due check from the collector's own start) runs fewer than 60.
  • Lite.Tests CollectorGridScheduleTests (build only here; CI runs it): mark at t0, then GetDueCollectorsForServer(id, t0+60 s) contains wait_stats and t0+59.9 s does not.
  • RED: with NextDue changed to now + interval, 4 of 34 pins fail (late run, stall, 1-minute simulation, Lite-shaped loop). Reverted, 34 of 34 pass. The Lite pin and the new 5-minute pin are not RED against the old code by design (the 5-minute pin asserts the grid's 12; the Lite pin needs the new overload, so on the old code it fails to compile).

Also run and green: StoreConfigProviderTests (14, 1 skipped, live), DocCommentHygieneTests, CommentFilterAdoptionTests, McpPayloadContractCensusTests, StorageCommandTimeoutTests, StartupCommandTimeoutTests. Build 0 errors, no CA/CS/IDE warnings. No live Postgres class is affected.

What to watch after the install

Body clustering: the "queued for a slot" info lines and the overrun diagnostic. Collectors now hold their grid, so more of them can be due in the same tick. Also the store's daily growth and WAL, and the sweep gate's busy share, which rise with the restored cadence (see above).

CHANGELOG

SECTION: Fixed
ENTRY: - Collectors run at their configured interval on large fleets, in Darling and Lite ([#4644]) - Each collector's next run was scheduled from when its pass started (Darling) or waited a full minute after the pass's work (Lite), so late starts accumulated: busy Darling stores ran 1-minute collectors about every 71 seconds and lost about 15% of per-minute samples. Both now keep a fixed grid and skip, rather than replay, a slot missed during a stall. Stores that were drifting collect correspondingly more data per day: up to about 20% more for the per-minute tables.
REF: [#4644]: #4644
IMPORTANT: After upgrading, 1-minute collectors run every 60 s instead of drifting to about 70 s on large fleets, so per-minute data and store growth rise by up to ~20% compared with 3.8.

…arling (#4640)

Lite's loop waited a full interval after each cycle's work and compared the
current time with each collector's own start. It now keeps a logical cycle
time on a fixed grid and records runs at that time. The grid helper moves to
the shared collectors project.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant