Repository navigation
Darling: file growth, long-running job, failed agent job, database state and forced plan alerts are tried again after every channel failed (#4752) - #4804
Merged
Conversation
…y after every channel failed (#4752) The tests for the file growth, long-running job, failed agent job, database state and forced plan alerts, the FailedSendBackoff unit tests and the reset-through-the-engine test. FailedSendBackoff is a stub that throws, so these fail against the engine as it is.
…ate and forced plan alerts are tried again after every channel failed (#4752) Adds FailedSendBackoff, which owns the failure streak the engine kept in a private dictionary: 1, 2, 4 minutes, capped at the cooldown, dropped by a delivery, and now also started over by a failure more than twice the cap after the last one. AlertEngine.AfterFire uses it. The five families that #4786 left out capture their delivery and call AfterFire, and each puts back the marker it wrote before delivery so the retry still counts as news: the per-file observation stamps, the saved failed-job watermark, the database-state announcement and the forced plan's observation. The long-running job needs none because its cooldown is per run.
erikdarlingdata
marked this pull request as ready for review
September 29, 2026 11:40
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #4752.
Why
An alert whose every channel failed (an HTTP 429 or 5xx, a timeout, an unreachable mail server) already retries after 1, 2, 4 ... minutes for nine SQL Server alert families instead of waiting out its cooldown. Five families still stamped their cooldown before delivery and stayed silent for the whole cooldown when nothing was delivered: Database File Growth, Long-Running Job, Failed Agent Job, Database State and Forced Plan Failing. Four of them also keep a second "already told" marker (written before delivery for file growth, failed job and forced plan; written after it, whatever the result, for database state), so a back-dated cooldown alone would have found nothing new on the retry sweep.
The failure streak had a gap of its own: only a delivery ended it, so a failure long after the condition cleared inherited the old count and started at the longer delay.
The failed-job marker is a saved watermark, and it was saved before delivery. When a fire that no channel received had no earlier watermark, the in-memory entry was removed but the saved row kept the new value. A restart inside the retry delay then read that failure as already announced, and the alert was lost for good.
What changes
FailedSendBackoff(PerformanceMonitor.Alerting/FailedSendBackoff.cs) owns the streak bookkeepingAlertEngine.AfterFiredid in a private dictionary:EveryChannelFailed(AlertDelivery?)(the rule, moved here),RecordFailure(family, key, nowUtc, cap)(counts a failure, returnsChannelFailureRetryDelay(failures, cap)) andRecordDelivered(family, key). It keeps the time of the last failure with the count, and a failure more than twicecapafter the previous one starts again at one minute. Twice, because a live streak's gap is a little over one cap. Thread-safe, in memory only.RecordFailure(..., out int failures)overload soAfterFirekeeps its log line, andTrackedCount. Streaks older than twice their own cap are dropped once the table passes 1024 pairs, which changes no answer (such a streak already restarts) and stops the per-run Long-Running Job keys from growing it forever.AfterFireuses oneFailedSendBackoffinstance. Its log line and back-dating are unchanged.ChannelFailureRetryDelaystays public static inAlertEngine. The privateEveryChannelFailednow calls the shared rule.var delivery = await FireAsync(...)and callAfterFire, and each makes sure the retry sweep still has something to fire on when every channel failed:IAlertStateStore.SaveFailedJobWatermarkAsyncsays the same. The change is in the shared engine, so Lite and Darling both get it; no store code changes.SaveDatabaseStateAlertedAsyncis not called when every channel failed, otherwise the retry of an edge-triggered state (a parked OFFLINE) reads as already announced. The comment above it now says a fire whose every channel failed is retried.Test plan
FailedSendBackoffthat throws. 17 of the 19 new test cases fail there (AlertEngineTests: 17 failed of 230). The other two are partial-failure cases (one channel delivered, no early retry) that pass either way by design.FailedJobs_tests failed there, and all 10 pass after the change:FailedJobs_EveryChannelFailedOnTheFirstFire_NeverSavesTheWatermark_AndTheRetryFiresInThisProcess, now asserts the saved state never received the failure's time (it used to assert the saved state kept it);FailedSendBackoffunit tests: 1, 2, 4 minutes then capped;RecordDeliveredstarts over; independent per family and key; a failure more than twice the cap later starts at one minute; one exactly twice the cap later continues; pruning; the delivery rule.Darling.TestsandLite.Testsbuild with 0 warnings.SaveFailedJobWatermarkAsyncorFailedSendBackoff, run again after the last edit (a comment-only change):AlertEngineTestsandDarlingSelfAlertTests, withDocCommentHygiene(559 tests, 0 failed, 3 skipped), andLiteAlertForwardingTestsandStoreRoundTripTests(43 tests, 0 failed).Darling.Tests(noDARLING_TEST_PG), run once on the change before that last comment-only edit: 17379 total, 0 failed, 1140 skipped (the live PostgreSQL classes), 1 not run. The count is the earlier 17375 plus the 4 new tests.Lite.Testswas last run on the earlier head of this branch (5607 total, 0 failed); after the change only the two Lite classes above were run.CHANGELOG
SECTION: Fixed
ENTRY:
REF:
[Darling: file growth, long-running job, failed agent job, database state and forced plan alerts are tried again after every channel failed (#4752) #4804]: Darling: file growth, long-running job, failed agent job, database state and forced plan alerts are tried again after every channel failed (#4752) #4804