Skip to content

Analysis facts, drill-downs and anomaly counts for blocking and deadlocks count events by when they happened, in both apps - #4913

Merged
erikdarlingdata merged 47 commits into
devfrom
fix/event-time-analysis-reads
Oct 1, 2026
Merged

erikdarlingdata merged 47 commits into
devfrom
fix/event-time-analysis-reads

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Builds on #4906 (an Azure SQL Database master's analysis findings skip separately monitored databases), which rewrote these same queries and is now merged; this PR's diff against dev is the window columns below.

What a user saw

The analysis findings, the drill-downs and the anomaly detector counted blocked-process reports (BPRs) and deadlocks by when they were COLLECTED. The grids and cards count by when the event HAPPENED (#4909). So a finding could cite a deadlock count for its window that the Deadlocks grid for the same window doesn't show.

What changed

Every analysis read of blocked_process_reports and deadlocks now windows, and buckets, on the event's own time.

Darling: each read is now the event column in the window, plus collection_time >= EventWindowFloor.For(start) (#3895), with no upper bound on collection_time.

pg_deadlocks (PostgreSQL targets): the exemplar count, the exemplar list and the MCP deadlock reader now agree. Each is COALESCE(occurred_at, collection_time) in the window, plus the floor. occurred_at is nullable, and a row without one counts by its collection time, as it did before. The MCP reader already used occurred_at but dropped a row without one; the fact and the drill-down used collection_time. The MCP reader also stamps, sorts and builds its report identity from the same expression, and the detail lookup matches on it, so a row without occurred_at is listed by its collection time (not as 0001-01-01 at the top of the page) and its identity resolves.

Lite: the event predicate alone. The reads:

Left alone, and why

  • Darling's baselines are continuous aggregates. A TimescaleDB continuous aggregate must bucket the hypertable's own time column, and these are hour-of-week baselines, where a seconds-long collection lag doesn't move a bucket. Changing them needs a rung and a rebuild.
  • The DMV blocking-snapshot arms: a snapshot's event_time IS its collection time.
  • The alert reads: a delivery cursor (see Blocking and deadlock windows count events by when they happened, in both apps' grids, MCP reads, slicers, trends and daily summary #4909).
  • The reads already on the event time: the blocking-chain fact and the reconstructed chains keep their event-time window and gain the collection_time floor, like every other read here.

Pins

  • Darling, EventTimeAnalysisReadsLiveTests: three rows per table:
    • (a) happened before the window and was collected inside it;
    • (b) happened and was collected inside it;
    • (c) happened inside it and was collected after it ended.
      The facts (both forms), the anomaly counts (both forms), the drill-downs and the pg_deadlocks reads count b and c, not a. The a/b/c row counts differ per table, so a collection-time window can't match by accident. For pg_deadlocks, a fourth row with no occurred_at counts by its collection time.
  • Lite, EventTimeAnalysisReadsTests: the same rows through the facts, the drill-downs, the anomaly counts and the baselines. An event at 10:58 collected at 11:03 lands in the 10:00 baseline bucket.
  • Existing pins updated: PgDeadlockRemaskTests binds the exemplar list's new floor parameter, PgTargetDeadlockDrillDownTests asserts the COALESCE form, and LocalClockBucketKeyTests reads the event column in the baseline's local-clock day. DarlingEventBaselineCoveredDaysTests still pins the same shared baseline body and collectors in both products, and now also pins the deliberate difference: Lite's Blocking and Deadlock arms wrap it with OnEventTime(..., "event_time") and (..., "deadlock_time"), and Darling's continuous-aggregate arms don't wrap.
  • Azure SQL Database: a monitored master's analysis findings no longer repeat the blocking and deadlocks of databases monitored as their own targets #4906's scoped variants moved too: TopDeadlocksSkippingSeparateSql, TopBlockingChainsSkippingSeparateSql, both DeadlockOutsideCountSql copies and BlockingSkippingSeparateCountSql. Lite's scoped deadlock count and its chain/fact scoping moved the same way. The reconstructed chains were already on the event time.
  • Parameter order: a plain read takes the floor as $4. A scoped read keeps Azure SQL Database: a monitored master's analysis findings no longer repeat the blocking and deadlocks of databases monitored as their own targets #4906's list at $4 and takes the floor as $5. Every caller binds the list first, then the floor.
  • Tighter pins (they passed before, and guard against a regression): the Darling drill-down seeds distinct waits and asserts the chains are {200, 300}; it also runs scoped (["GP"] over HS rows), which exercises Azure SQL Database: a monitored master's analysis findings no longer repeat the blocking and deadlocks of databases monitored as their own targets #4906's skipping variants. Lite adds scoped cases for the deadlock fact and the top deadlocks.
  • The two newer pins, RED at their tests commit on a local live store: the NULL-occurred_at reader pin (the stamp, the sort and the identity round trip), and BlockingChainReadsFloorTests (2 of 2).
  • Seeders that set no event time now set it equal to the collection time (test data only): EventBaselineCoveredDaysTests, Lite's AnomalyTileWindowEndTests, and AzureMasterAnalysisScopeTests.SeedBprAsync.
  • Darling's AnomalyTileWindowEndTests text pin counts the open-ended window on event_time < $3 / deadlock_time < $3 as well, and still rejects the closed spellings.
  • RED before the fix, Darling: a local live store at the tests-first head 3653d8efb. EventTimeAnalysisReadsLiveTests failed 6 of 6, on assertions:
    • BlockingAndDeadlockFacts_… (plain and Azure-master forms);
    • AnomalyDetectorCurrentWindowCounts_… (both forms);
    • DrillDownTopDeadlocksAndTopChains_…;
    • PgDeadlockCapture_CountsByOccurrenceWithTheCollectionTimeFallback_….
      CI's RED run was skipped while Actions is degraded.
  • RED before the fix, Lite: reasoned from the old query text, which windowed on collection_time. Every pin seeds row (a) (happened before the window, collected inside it) and row (c) (happened inside it, collected after it ended). The old predicate counts (a) and drops (c).
    • BlockingEventsFact_… and DeadlocksFact_… seed a second (c) row and assert a count of 3 (b + 2c). The old read counts a + b = 2.
    • TopDeadlocks_… asserts victims [c, b], and TopBlockingChains_… waits [300, 200]. The old read returns a and b.
    • AnomalyDetector_CurrentCounts_… asserts 6 events (3 b + 3 c). The old read counts 1 a + 3 b = 4.
    • BaselineBuckets_CountEventsByWhenTheyHappened (blocking and deadlocks) asserts that the 10:58-event / 11:03-collected row lands in hour 10 and not 11. The old read buckets by collection, into hour 11.
  • At the fix, after the Azure SQL Database: a monitored master's analysis findings no longer repeat the blocking and deadlocks of databases monitored as their own targets #4906 merge, locally: 23 Darling classes green on a live store, 0 failed. They are the new class, AzureMasterAnalysisScopeLiveTests, AnomalyDetectorErrorGuardLiveTests, AnalysisCoverageLivePostgresTests, AnalysisFactsReadRunsDetectorParityTests, CountFamilyZeroHistoryTests, DarlingAnomalyBaselineTests, PgTarget*, PgDeadlockRemaskTests, PgFactCollectorTests, McpPageContractTests, DarlingAnalysisPipelineTests, AnomalyTileWindowEndTests, AzureMasterScopeTests and Azure SQL Database: a monitored master's analysis findings no longer repeat the blocking and deadlocks of databases monitored as their own targets #4906's new classes. Lite.Tests builds with 0 errors; CI runs it.

CHANGELOG

The collection-time windows shipped in v3.8.0, so this is a Fixed entry.

SECTION: Fixed
ENTRY: - Analysis findings, drill-downs and anomaly checks count blocking and deadlocks by when they happened, matching the grids ([#4913]) - These reads counted events by when they were collected, so a finding's count for a window could differ from the grid's for the same window. On PostgreSQL targets, the deadlock fact, drill-down and MCP reader now also agree, and a deadlock without a timestamp is counted by when it was collected.
REF: [#4913]: #4913

… and deadlocks that belong to databases monitored as their own targets
…s for databases monitored as their own targets
…ases, and the deadlock every-process rule lives in one place
…he sustained-blocking template warns about a master target
…findings skip databases monitored as their own targets
…indings skip databases monitored as their own targets
…ck findings skip databases monitored as their own targets
…k findings skip databases monitored as their own targets
…its closing parenthesis, since another argument now follows it
…ence and fact reads

BLOCKING_CHAIN and the blocking/deadlock drill-downs skip databases monitored as their own targets.
The fact and compare reads fill the list like AnalyzeAsync. The deadlock count counts a deadlock whose
victim database is outside the list in SQL and parses graphs only for the rest. {SCOPE} sits at the end
of its line so an empty list gives the old text. The static provider is internal and a throwing provider
falls back to unscoped.
… master target; bound the deadlock graph read

BLOCKING_CHAIN drops pairs of databases monitored as their own targets. The top-deadlock and top-blocking drill-downs apply the same rule. Deadlocks whose named victim database is not separately monitored are counted in SQL and their graphs are not read; the anomaly path uses the detector's command timeout. The list is passed raw and both sides fold with one lower(). MCP analyze_server and the web read tools resolve the list per call from the live registry.
… and recognises Azure SQL Database by its stored engine edition
…ostgreSQL deadlock capture count events by when they happened, not when they were collected
…d row is checked by its graph, and the unscoped text is unchanged
…a failing scope resolver degrades to unscoped
…downs, including the Azure master scoped variants, window blocked-process reports and deadlocks on when the event happened, and the PostgreSQL deadlock capture on occurred_at with a collection_time fallback
…, including the Azure master scoped variants, count blocked-process reports and deadlocks by when they happened, and the seeders that set no event time now set it equal to the collection time
…ected event, so the event-time count differs from the collection-time one
…stamp pg deadlock read pin, and floor pin for the two chain reads
… MCP deadlock reader stamps, sorts and identifies a NULL-occurred_at report by its collection time
…s a pinned, deliberate difference from Darling's continuous aggregate
# Conflicts:
#	Darling/PerformanceMonitor.Darling.Analysis/PgAnomalyDetector.cs
#	Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Blocking.cs
#	Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Waits.cs
#	Lite.Tests/AzureMasterAnalysisScopeTests.cs
#	Lite/Analysis/AnomalyDetector.cs
#	Lite/Analysis/DrillDownCollector.Blocking.cs
#	Lite/Analysis/DuckDbFactCollector.Waits.cs
#	Lite/Analysis/SeparatelyMonitoredScope.cs
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 1, 2026 18:21
@erikdarlingdata
erikdarlingdata merged commit 9d61ec4 into dev Oct 1, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/event-time-analysis-reads branch October 1, 2026 18:22
erikdarlingdata added a commit that referenced this pull request Oct 1, 2026
Dev moved the blocking reads to event_time (#4909, #4913). A read that now
windows on event_time calls StoredEventCopies without collectedFrom and
keeps its event_time bounds in its filter: a stored copy keeps its first
copy's event_time, so both are inside the window or both are outside it.
The alert engine's read of blocked process reports still windows on
collection_time, so it keeps collectedFrom.

OnEventTime now swaps only the events CTE's own local-clock keys and
window. Swapping every collection_time in the CTE would also rewrite the
copy rule inside the blocking baseline's source, and that rule must stay
on collection_time.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant