Repository navigation
Keep recurring collectors on a fixed grid instead of drifting from the sweep body start - #4644
Merged
Merged
Conversation
…arling (#4640) Lite's loop waited a full interval after each cycle's work and compared the current time with each collector's own start. It now keeps a logical cycle time on a fixed grid and records runs at that time. The grid helper moves to the shared collectors project.
erikdarlingdata
marked this pull request as ready for review
September 28, 2026 20:13
This was referenced Sep 28, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #4636, Fixes #4640
What changed
The sweep set each due collector's next due time to
now + interval, wherenowis when the server's sweep body started. Every late start pushed the schedule back, so busy stores collected about every 71 seconds instead of 60. A new pure helper in the shared collectors project,CollectorCadence.NextDue(due, now, interval)(used by Darling and Lite), returns the previous due time plus the interval. If that slot is already past, it skips the missed slots and returns the next grid slot, so a stalled server resumes on its grid with no catch-up burst.RunDueCollectorsAsyncuses it. Comments on the advance and onComputeSeededNextDuenow say the steady-state advance is grid-based.Interval changes: the schedule reload re-seeds or caps
NextDuewhen the effective interval changes (DarlingWorker.cs~4692–4709), and the helper reads the interval per call, so a change takes effect on the next advance.Lite (#4640): the loop waited a full minute after each cycle's work, and the due check compared the current time with the collector's own start time, so a 1-minute collector missed cycles whenever a cycle ran long.
CollectionBackgroundServicenow keeps a logical cycle time on a fixed grid (CollectorCadence.NextDue, waiting only the remainder to the next slot).RunDueCollectorsAsync(cycleStartUtc, ...)passes that time to a newScheduleManager.GetDueCollectorsForServer(serverId, atUtc)overload and records each run at the cycle time (RunCollectorAsync(server, name, scheduledAtUtc, ct)->MarkCollectorRunForServer). The old overloads keep their meaning (they use the current time). The due check now compares exact grid times: a 1-minute collector is due every cycle, a 5-minute one every fifth.Field read on one production store before the change: 1-minute collectors had p50 71.5-71.9 s between runs (~50 runs/server/hour); 5-minute collectors p50 ~307 s (10.4 runs/hour).
The volume effect. Restoring the configured cadence means more collection. On that store, 1-minute collectors go from ~50 to 60 runs per server per hour (+20%) and 5-minute collectors from ~10.4 to 12 (+15%), so the per-minute tables take ~15–20% more rows and store bytes per day, and the store's WAL rises with them. The sweep gate's busy share rises in step: on that store from 49% to ~59% at width 6.
A manual "collect now" run in Lite (
RunCollectorAsyncwithout a cycle time) still records its own start time. After one, that collector's next due check compares the next cycle time with a wall-clock time, so it can skip one cycle. That's the same as today, and only after a manual run.Pins (
DarlingSweepSchedulingTests,Lite.Tests/CollectorGridScheduleTests)now + intervalrule: 53 runs. The pin asserts 59–61 and that the old rule is lower.CollectorGridScheduleTests(build only here; CI runs it): mark at t0, thenGetDueCollectorsForServer(id, t0+60 s)containswait_statsandt0+59.9 sdoes not.NextDuechanged tonow + interval, 4 of 34 pins fail (late run, stall, 1-minute simulation, Lite-shaped loop). Reverted, 34 of 34 pass. The Lite pin and the new 5-minute pin are not RED against the old code by design (the 5-minute pin asserts the grid's 12; the Lite pin needs the new overload, so on the old code it fails to compile).Also run and green:
StoreConfigProviderTests(14, 1 skipped, live),DocCommentHygieneTests,CommentFilterAdoptionTests,McpPayloadContractCensusTests,StorageCommandTimeoutTests,StartupCommandTimeoutTests. Build 0 errors, no CA/CS/IDE warnings. No live Postgres class is affected.What to watch after the install
Body clustering: the "queued for a slot" info lines and the overrun diagnostic. Collectors now hold their grid, so more of them can be due in the same tick. Also the store's daily growth and WAL, and the sweep gate's busy share, which rise with the restored cadence (see above).
CHANGELOG
SECTION: Fixed
ENTRY: - Collectors run at their configured interval on large fleets, in Darling and Lite ([#4644]) - Each collector's next run was scheduled from when its pass started (Darling) or waited a full minute after the pass's work (Lite), so late starts accumulated: busy Darling stores ran 1-minute collectors about every 71 seconds and lost about 15% of per-minute samples. Both now keep a fixed grid and skip, rather than replay, a slot missed during a stall. Stores that were drifting collect correspondingly more data per day: up to about 20% more for the per-minute tables.
REF: [#4644]: #4644
IMPORTANT: After upgrading, 1-minute collectors run every 60 s instead of drifting to about 70 s on large fleets, so per-minute data and store growth rise by up to ~20% compared with 3.8.