Skip to content

The alert notebook misses a resolution that lands more than 24 hours after the alert #4755

Description

@erikdarlingdata

The alert notebook shows a status for each alert. It checks three things in order. First, a resolution row after the alert. Next, a later firing of the same metric. Last, whether the metric's collector ran since the alert.

The first two checks read only config_alert_log rows from the alert time to 24 hours after it, at most 200 rows, newest first. So the notebook misses a resolution in three cases:

  • A store self-alert that clears more than 24 hours after it fired. Its Cleared row is outside the window. Each store self-alert maps to a Cleared alias (DarlingTriageEndpoint.cs:126-139), so the resolution row exists. The page reads "Fired again", or for a fleet-level store alert, "Unknown (not collected since ...)".
  • Fleet-level metrics carry no server id, so one 200-row limit covers every server. Newest first drops the oldest rows, and that is where a resolution sits. Nobody measured whether a real fleet writes more than 200 rows in 24 hours.
  • The read skips dismissed rows, and the viewer's "Dismiss all" marks Cleared rows too. So dismissing the alert list also hides the resolution.

This affects the display only. The page never shows a false "Resolved" or a false "ongoing".

Where (as of 64cd886)

Paths are under Darling/PerformanceMonitor.Darling.Service/ unless named in full.

  • The status checks: AlertNotebookEndpoint.cs:911-961, and the history checks at :1065-1151.
  • The window and the limit: AlertNotebookEndpoint.cs:170-179, DarlingTriageEndpoint.cs:527 and Mcp/DarlingAlertReader.cs:98-103.
  • No server id for fleet-level metrics: AlertNotebookEndpoint.cs:85-90.
  • Dismissed rows are skipped: Mcp/DarlingAlertReader.cs:133-134. "Dismiss all": Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertHistory.cs:229-234.

Fix

Add one targeted read for the first resolution row after the alert, with no 24-hour limit and no shared row limit.

  1. It matches the alias and recovery names (DarlingTriageEndpoint.ResolutionAliases, and NotebookRecoveryEdges at AlertNotebookEndpoint.cs:1046), not the firing metric.
  2. It includes dismissed rows.
  3. It filters by server_id when it is known. The index on config_alert_log covers that (Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs:2695).

Do the same for the second check, the later firing of the same metric.

Test that must fail first

  • A store self-alert whose Cleared row lands 30 hours later reads "Resolved at T".
  • So does one whose Cleared row was dismissed.
  • With 250 newer rows from other servers inside the window, the resolution is still found.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingin-progressActively being worked by a local session or its agents (PR open or in flight)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions