The alert engine stamps its cooldown before delivery. For deadlocks and blocking, it also moves and saves the count watermark when the alert fires. When every channel then fails with a transient error (an HTTP 429 or 5xx, a timeout, or DNS), nothing retries. Only the two daily documents retry on ChannelFailed.
So a standing condition is lost for the cooldown, which is 5 minutes by default. A deadlock or blocking alert is not announced again until the count rises. The row is recorded as failed with send_error, so the history shows it.
Webhook posts are four awaits in a row, each of up to 30 seconds, with no cancellation token.
Both apps run this engine. Lite's deliverer returns no result, and the engine reads that as delivered. So the Lite half of the fix needs Lite to report a failed send too.
- The stamp before delivery:
PerformanceMonitor.Alerting/AlertEngine.cs:781, :1094, :1428, :1540 and :1661. Delivery is at :3086-3097.
- The watermark at fire:
PerformanceMonitor.Alerting/RollingCountAlertGate.cs:85-88, and AlertEngine.cs:1067-1071 and :1517-1521.
- Only the daily documents retry:
PerformanceMonitor.Alerting/IAlertDeliverer.cs:100-119.
- Lite returns no result:
Lite/Services/LiteAlertDeliverer.cs:190-194.
- The default cooldown:
Darling/PerformanceMonitor.Darling.Service/DarlingConfig.cs:849.
- The posts:
PerformanceMonitor.Notifications/WebhookAlertService.cs:312-337. The 30-second client is at :2591-2593, and the send without a token is at :2634. The caller is PerformanceMonitor.Notifications/EmailSendCore.cs:265.
Fix
Reuse the result the daily documents already use (DeliverAndReportAsync), so a ChannelFailed result skips the engine's stamp and the watermark move. First add a timeout, cancellation and backoff to the posts. Without them, one dead webhook stalls each sweep for 30 seconds per retry.
Test that must fail first
Every channel fails with a transient error. The next evaluation fires again, and a deadlock alert keeps its watermark. Cover both apps.
The alert engine stamps its cooldown before delivery. For deadlocks and blocking, it also moves and saves the count watermark when the alert fires. When every channel then fails with a transient error (an HTTP 429 or 5xx, a timeout, or DNS), nothing retries. Only the two daily documents retry on
ChannelFailed.So a standing condition is lost for the cooldown, which is 5 minutes by default. A deadlock or blocking alert is not announced again until the count rises. The row is recorded as failed with
send_error, so the history shows it.Webhook posts are four awaits in a row, each of up to 30 seconds, with no cancellation token.
Both apps run this engine. Lite's deliverer returns no result, and the engine reads that as delivered. So the Lite half of the fix needs Lite to report a failed send too.
Where (as of 64cd886)
PerformanceMonitor.Alerting/AlertEngine.cs:781,:1094,:1428,:1540and:1661. Delivery is at:3086-3097.PerformanceMonitor.Alerting/RollingCountAlertGate.cs:85-88, andAlertEngine.cs:1067-1071and:1517-1521.PerformanceMonitor.Alerting/IAlertDeliverer.cs:100-119.Lite/Services/LiteAlertDeliverer.cs:190-194.Darling/PerformanceMonitor.Darling.Service/DarlingConfig.cs:849.PerformanceMonitor.Notifications/WebhookAlertService.cs:312-337. The 30-second client is at:2591-2593, and the send without a token is at:2634. The caller isPerformanceMonitor.Notifications/EmailSendCore.cs:265.Fix
Reuse the result the daily documents already use (
DeliverAndReportAsync), so aChannelFailedresult skips the engine's stamp and the watermark move. First add a timeout, cancellation and backoff to the posts. Without them, one dead webhook stalls each sweep for 30 seconds per retry.Test that must fail first
Every channel fails with a transient error. The next evaluation fires again, and a deadlock alert keeps its watermark. Cover both apps.