Skip to content

Fleet overview, daily summary and server summary: an Azure master counts only its own blocking and deadlocks, and its lists say where its databases' events are counted - #4925

Merged
erikdarlingdata merged 22 commits into
devfrom
fix/fleet-azure-master-scope
Oct 2, 2026
Merged

erikdarlingdata merged 22 commits into
devfrom
fix/fleet-azure-master-scope

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What a user saw

On a fleet with an Azure SQL Database logical server registered at master and its databases registered on their own, Darling counted the databases' blocking events and deadlocks twice. The Fleet Overview header showed Blocking 96 (master 48 + GP 25 + HS 23) and Deadlocks 91 (46 + 23 + 22), and master sat in "Needs attention" as Critical on those same events. Master's daily summary, the Performance Calendar, the fleet sweep's band, MCP get_server_summary and the Darling Viewer's fleet totals and per-server summary did the same. Master's Blocking, Deadlocks and Locking lists showed those rows with nothing saying why the card disagreed.

Why

A master registration reads server-wide blocked-process reports and deadlocks, which include the separately monitored databases' events. Alerts and analysis already skip those databases' events on a master target. These reads didn't: they grouped every stored event by server.

What changed

Every read below takes the separately monitored list analysis uses (AzureMasterScope.SeparatelyMonitoredDatabases). It applies only to an Azure SQL Database master: the server's registration (servers.sql_engine_edition, stamped on every connect) AND its newest stored server_properties.engine_edition must both be 5, on every surface. The service's stored-edition read (DarlingWorker.StoredEngineEditionSql) now tests both, so a pair that disagrees reads unscoped everywhere and #4906's scoped SQL (PgFactCollector.BlockingSqlSkippingSeparate, CountDeadlocksSkippingSeparateAsync: a deadlock counts on master unless every process in its graph is in the list). The DMV blocking arm is unchanged (master's DMV sees only its own sessions).

  • The fleet card (DarlingFleetReader.GetFleetOverviewAsync: /api/fleet, the web read mirror, MCP get_fleet_overview): master's blocked-process count and max wait and its deadlock count come from the scoped reads. The cards, header totals, bands and "Needs attention" follow.
  • MCP get_server_summary (DarlingHealthReader.GetServerSummaryAsync): the same scoped reads over its one-hour window.
  • The daily summary (DailySummarySql.RangeSql via DailySummaryAzureMasterScope: MCP get_daily_summary / get_daily_summary_range, the read mirror, the fleet sweep's band through GetWindowSignalsAsync, and the Viewer's Performance Calendar; a master that is disconnected when the sweep fires gets its list from the stored edition and the registry, as MCP does): the deadlock and blocked-process steps count with a FILTER, and master's graph-carrying deadlocks are counted per day with the every-process rule. The windows, day bucketing and DMV arm are byte-identical. Those steps read the raw tables on every retention tier, so every retained day is scoped. The range cache key carries the list. A day whose only events are siblings' is still a present day, with a count of 0.
  • The lists keep their rows (master captures server-wide), with one line above master's Blocking and Deadlocks lists and its Locking & Contention panel, in the Viewer and on the web server page: "Events from databases monitored as their own servers are listed here and counted under those servers." MCP get_blocking and get_deadlocks return it as separately_monitored_note, with the list as separately_monitored_databases; get_object_locking returns the note. The note shows only when the list is non-empty.
  • The Darling Viewer computes the same list from the store (the monitored-server registrations and the same two-edition rule). It reads only those columns, never a credential column. It scopes the per-server summary and the calendar, and subtracts the over-count from the fleet totals so they agree with the cards.
  • When something fails: a failed list lookup, or a failed scoped read, keeps that master's unscoped counts and logs a warning. The rest of the fleet read succeeds. A failed note lookup shows no note, and the rows still return. A scoped read that times out is followed by the unscoped read, so one call can take up to about twice its read deadline.
  • get_server_summary's scoped read ends a day ahead of now, as the Viewer's does, so a row stamped ahead by clock skew counts the same in both.
  • Unchanged: every other server, and a master with no separately monitored databases (the same SQL as before); the newest-deadlock time; the fleet SQL constants.
  • Cost: for a non-master, the Viewer adds one indexed edition check per card, calendar, Blocking tab or Locking panel refresh. For a master, one list read plus one graph read bounded by the window. A range-cache hit re-reads only the open day.

Pins

  • FleetOverviewAzureMasterScopeLiveTests: master + GP + HS on one host: master's card shows Blocking 3 and Deadlocks 1 and the header counts each event once; a sibling's longer wait leaves master's max wait at master's own; a non-Azure server and a sibling-less master are unchanged; with no resolver, a resolver that throws, or a scoped read that throws, the old counts and every card; get_server_summary matches the card, and an edition pair that disagrees either way round reads unscoped.
  • ViewerFleetAzureMasterScopeLiveTests: the same arrange through the Viewer's summary and fleet totals, the two-edition rule, a failed scoped read falling back, and the registry read only for masters and once per fleet refresh.
  • DailySummaryAzureMasterScopeTests, DailySummaryAzureMasterScopeLiveTests, ViewerDailySummaryAzureMasterScopeLiveTests: the scoped SQL keeps the windows and the DMV arm, every tier is scoped, the cache key carries the list, master's days exclude sibling events, the service and the Viewer agree, and a throwing scoped read falls back.
  • SeparatelyMonitoredListNoteTests: the sentence, and when it shows on each surface.
  • AzureMasterAnalysisScopeLiveTests: the stored-edition rule (including a registration of 8 with a newest stored edition of 5), and a source pin that the sweep scopes a master with no runtime.
  • ServerPageTabsTests: table()'s full signature, now with moreNoteKeys.
  • The fleet pins fail on assertions at the tests commit (master blocking expected 3, actual 6), and pass with the fix.

Tests run

  • Builds: Lite.Tests and Darling.Tests, 0 warnings and 0 errors.
  • Darling.Tests in-process, on a live PostgreSQL store, all passing: FleetOverviewAzureMasterScopeLiveTests (9), AzureMasterAnalysisScopeLiveTests (21), DailySummaryAzureMasterScopeLiveTests (4), ViewerDailySummaryAzureMasterScopeLiveTests (3), ViewerFleetAzureMasterScopeLiveTests (8), DailySummaryAzureMasterScopeTests (6), SeparatelyMonitoredListNoteTests (7), ServerPageTabsTests (19), PgIndexBloatCoverageTests (16), FleetSweepSignalReadLivePostgresTests, FleetSweepCadenceKnobRungTests, DarlingMcpHealthToolsLivePostgresTests, DarlingSelfAlertTests, CollectorMemoryKnobTests, and the T-SQL, doc-comment, repo-file and deadline-scanner guards.
  • The new edition and no-runtime pins fail with the worker change reverted, and pass with it.

CHANGELOG

None as its own line: this widens the existing Fixed entry for the fleet overview's Azure master count (#4894) to name the daily summaries, calendars, the sweep band, get_server_summary and the list notes. The double count was present in v3.8.0.

…gain for its own databases

The reader gains the optional resolver parameter (not yet used), so the pins compile and fail on their assertions.
…; one engine-edition rule; the registry is read only for masters
…et_server_summary counts an Azure master once
…tored databases' events are counted under those servers
…bases' rows stay; the note's lookup fails safe
…ine-edition rule

The fleet sweep read a master with no runtime unscoped, so its earlier events counted twice in the sweep band; it now asks the stored edition and the registry. StoredEngineEditionSql also requires servers.sql_engine_edition = 5, because the connect probe and the server_properties collector are separate writers and can disagree. The server summary's scoped read ends a day ahead like the Viewer's, the scope fallbacks log the exception, and table()'s doc names moreNoteKeys.
@erikdarlingdata erikdarlingdata changed the title Fleet overview: an Azure master's card counts only its own blocking and deadlocks, not its databases' again Fleet overview, daily summary and server summary: an Azure master counts only its own blocking and deadlocks, and its lists say where its databases' events are counted Oct 2, 2026
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 2, 2026 00:25
@erikdarlingdata
erikdarlingdata merged commit b60510b into dev Oct 2, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/fleet-azure-master-scope branch October 2, 2026 00:26
erikdarlingdata added a commit that referenced this pull request Oct 2, 2026
…wn databases, and the Viewer card shows no sibling's Last (#4932)

#4925 scoped an Azure SQL Database master's blocking and deadlock counts to its own databases but left the last-seen times alone, so a master whose own databases had no deadlocks showed "0 deadlocks, last seen minutes ago" on the fleet card from a separately monitored database's deadlock, and the Darling Viewer's Overview card read "Last: N ago" from a sibling's event under 0/0. The fleet card's deadlock_last_seen now comes from the same single pass as the scoped count, by the count's own rule, so the two can't disagree and a failure keeps that master's whole unscoped row; get_server_summary makes one graph pass instead of two. In the Viewer, a master with separately monitored databases shows no "Last" on its Overview card for blocking or deadlocks, because an own-database newest over all stored history would need an all-history read or a cache on every refresh; its counts stay scoped and its lists keep every row with their note. Every other server keeps its "Last".
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant