The alert notebook shows a status for each alert. It checks three things in order. First, a resolution row after the alert. Next, a later firing of the same metric. Last, whether the metric's collector ran since the alert.
The first two checks read only config_alert_log rows from the alert time to 24 hours after it, at most 200 rows, newest first. So the notebook misses a resolution in three cases:
- A store self-alert that clears more than 24 hours after it fired. Its Cleared row is outside the window. Each store self-alert maps to a Cleared alias (
DarlingTriageEndpoint.cs:126-139), so the resolution row exists. The page reads "Fired again", or for a fleet-level store alert, "Unknown (not collected since ...)".
- Fleet-level metrics carry no server id, so one 200-row limit covers every server. Newest first drops the oldest rows, and that is where a resolution sits. Nobody measured whether a real fleet writes more than 200 rows in 24 hours.
- The read skips dismissed rows, and the viewer's "Dismiss all" marks Cleared rows too. So dismissing the alert list also hides the resolution.
This affects the display only. The page never shows a false "Resolved" or a false "ongoing".
Paths are under Darling/PerformanceMonitor.Darling.Service/ unless named in full.
- The status checks:
AlertNotebookEndpoint.cs:911-961, and the history checks at :1065-1151.
- The window and the limit:
AlertNotebookEndpoint.cs:170-179, DarlingTriageEndpoint.cs:527 and Mcp/DarlingAlertReader.cs:98-103.
- No server id for fleet-level metrics:
AlertNotebookEndpoint.cs:85-90.
- Dismissed rows are skipped:
Mcp/DarlingAlertReader.cs:133-134. "Dismiss all": Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertHistory.cs:229-234.
Fix
Add one targeted read for the first resolution row after the alert, with no 24-hour limit and no shared row limit.
- It matches the alias and recovery names (
DarlingTriageEndpoint.ResolutionAliases, and NotebookRecoveryEdges at AlertNotebookEndpoint.cs:1046), not the firing metric.
- It includes dismissed rows.
- It filters by
server_id when it is known. The index on config_alert_log covers that (Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs:2695).
Do the same for the second check, the later firing of the same metric.
Test that must fail first
- A store self-alert whose Cleared row lands 30 hours later reads "Resolved at T".
- So does one whose Cleared row was dismissed.
- With 250 newer rows from other servers inside the window, the resolution is still found.
The alert notebook shows a status for each alert. It checks three things in order. First, a resolution row after the alert. Next, a later firing of the same metric. Last, whether the metric's collector ran since the alert.
The first two checks read only
config_alert_logrows from the alert time to 24 hours after it, at most 200 rows, newest first. So the notebook misses a resolution in three cases:DarlingTriageEndpoint.cs:126-139), so the resolution row exists. The page reads "Fired again", or for a fleet-level store alert, "Unknown (not collected since ...)".This affects the display only. The page never shows a false "Resolved" or a false "ongoing".
Where (as of 64cd886)
Paths are under
Darling/PerformanceMonitor.Darling.Service/unless named in full.AlertNotebookEndpoint.cs:911-961, and the history checks at:1065-1151.AlertNotebookEndpoint.cs:170-179,DarlingTriageEndpoint.cs:527andMcp/DarlingAlertReader.cs:98-103.AlertNotebookEndpoint.cs:85-90.Mcp/DarlingAlertReader.cs:133-134. "Dismiss all":Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertHistory.cs:229-234.Fix
Add one targeted read for the first resolution row after the alert, with no 24-hour limit and no shared row limit.
DarlingTriageEndpoint.ResolutionAliases, andNotebookRecoveryEdgesatAlertNotebookEndpoint.cs:1046), not the firing metric.server_idwhen it is known. The index onconfig_alert_logcovers that (Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs:2695).Do the same for the second check, the later firing of the same metric.
Test that must fail first