Skip to content

Lite counts each blocked process report, long query completion and system_health event once after a reset stored it twice - #4918

Merged
erikdarlingdata merged 18 commits into
devfrom
fix/lite-archive-view-dedupe-after-reset
Oct 1, 2026
Merged

erikdarlingdata merged 18 commits into
devfrom
fix/lite-archive-view-dedupe-after-reset

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What users see

Before #4887, the first collection after Lite's 512 MB archive-and-reset read the collectors' 10-minute fallback window again. It stored blocked process reports, long query completions and system_health events that the archive already held. #4887 stopped new copies, but the copies already stored stay in the archive files.

Every read of these three tables counted such an event twice when both copies fell in its window. These reads serve:

  • blocking counts and alerts
  • the blocking grids, chains, trend, slicer and severity chart
  • the daily summary and calendar
  • analysis facts, anomaly counts and baselines
  • the long query grid and MCP tool
  • the system_health parsers

A watermark read that fails also takes the fallback window, so it can still store a copy.

In the 30-day reads measured below, 7 of 942 blocked process reports and 9 of 5,189 system_health events of one type were copies.

Cause

  • PerformanceMonitor.Collectors/BlockedProcessReportCollector.cs:759, LongQueryCompletionsCollector.cs:230 and SystemHealthEventsCollector.cs:159: a collector with no watermark reads back the fallback window.
  • Before Lite keeps its watermarks and alert state across the 512 MB archive-and-reset, so it stops storing events twice #4887, the watermark came from the live table alone, which ArchiveService.ArchiveAllAndResetAsync empties. So the first cycle after each reset took the fallback window and stored those events again.
  • Lite/Services/RemoteCollectorService.cs:1480, :1579, :1659 and :1721: a watermark read that fails also falls back. These catches logged nothing.
  • Lite/Database/DuckDbInitializer.cs: each archive view for these three tables is a plain union of the live table and the archive files. So every read saw both copies.

Changes

  • Lite/Database/StoredEventCopies.cs (new) is the one place that reads v_blocked_process_reports, v_long_query_completions and v_system_health_events. Each read passes its own filter, and the helper drops the copies that a later batch stored, after that filter.
  • The rule: for each identity, every row of the first batch that stored it stays, and the copies from later batches go. That is DENSE_RANK() OVER (PARTITION BY identity ORDER BY collection_time) = 1. The helper groups the rows on the identity alone to find each identity's MIN(collection_time). It joins that back to the rows with IS NOT DISTINCT FROM on every part of the identity, so a NULL part matches a NULL part.
  • Rows from one batch are never collapsed. A collector stamps one collection_time on a whole run, strictly later than its previous run, so identical rows from one batch share it.
  • The identity of a blocked process report or a system_health event is server_id, event_time and a 64-bit hash() of the event's XML. A long query completion has no XML. Its identity is the server, database, event time, statement text, session and event sequence, all compared exactly.
  • A row with no event time or no text is never collapsed. Two more parts add that row's own id and collection_time to its key. They test the raw text, not its hash, because hash(NULL) is not NULL.
  • XmlKey in StoredEventCopies.cs is the only place the hash is named. Two different reports, or system_health events, that share a server and event time merge only if their XML hashes collide. The chance is about 2^-64 per pair. A collision only affects reads: it hides one row from one read and never deletes anything. Nothing stores the hash.
  • Why the rule is grouped and hashed: a test ran 94,200 reports at Lite's 1 GB memory_limit. There, a window over the identity ran out of memory. It did so with the XML compared exactly and with the XML keyed by md5_number. A window carries every row, XML included, through the operator, while the grouped side holds only the keys. md5_number gives a 128-bit key. On DuckDB 1.5.5, though, it read the XML at about 140 MB/s on one thread, 11 times slower than hash(). The numbers are under Performance.
  • Most reads now filter on event_time, and they need no look-back. A copy keeps its first copy's event_time, so both are inside the filter or both are outside it.
  • Three calls still filter on collection_time: the two long query reads and the alert engine's form of GetRecentBlockedProcessReportsAsync. Each passes its lower bound to the helper. The rule then also reads CollectorContext.EventFallbackWindow before that bound. So a copy stored inside the window, whose first copy was stored just before the window, is dropped too, and the first copy stays outside.
  • That look-back assumes the monitored server's clock, which stamps event_time, is not ahead of the collector's. A server clock that is ahead can put a first copy further back. A copy stored just inside the window's start is then read.
  • PerformanceMonitor.Collectors/CollectorContext.cs:70 adds EventFallbackWindow (10 minutes). The three collectors and the helper use it, and no copy of the literal is left.
  • The archive views stay a plain union. A rule in the view runs over the whole archive on every read. collection_time is not part of an event's identity, so a read's collection_time filter cannot run below that rule. The comment in DuckDbInitializer.cs that says so is brought up to date.
  • Lite/Analysis/BaselineProvider.cs: OnEventTime moves the blocking baseline's events CTE to event_time. It swapped every collection_time in that CTE, which also rewrote the helper's rule inside it, and the rule then kept every copy. It now swaps only the CTE's own keys and window.
  • The four watermark-read catches in RemoteCollectorService.cs now log a WARN that names the read, the table, the server id and the exception type.
  • The Azure master scope filter ({SCOPE}) on three blocking reads sits inside the helper's filter. The helper repeats that filter in both of its scans, and one Replace covers the whole statement.
  • Deadlocks: handled at the view level in Azure SQL Database: a master registration stores each deadlock once, cleans up the copies already stored, and loses no late or ring-only event #4910.
  • A bare read of the three tables themselves is caught by ArchivableTableBareReadSweepTests in Lite advice, High Impact Queries and the Agent check read archived rows after the 512 MB archive step #4921, because all three are in ArchiveService.ArchivableTables.

Every read of the three views

There are 23 reads. 20 go through the helper, in 21 calls, because read 10 has two forms. The other 3 read the view directly.

# Site Read Through Serves
1 Lite/Analysis/AnomalyDetector.cs:790 current_blocking count, by event_time helper anomaly counts
2 Lite/Analysis/BaselineProvider.cs:717 blocking baseline events, by event_time helper baselines
3 Lite/Analysis/BaselineProvider.cs:775 blocking per minute, by event_time helper baselines
4 Lite/Analysis/DrillDownCollector.Blocking.cs:92 top blocking chains, by event_time helper drill-downs
5 Lite/Analysis/DrillDownCollector.Blocking.cs:157 reconstructed chains, by event_time helper drill-downs
6 Lite/Analysis/DuckDbFactCollector.Waits.cs:254 blocking facts, by event_time helper analysis facts
7 Lite/Analysis/DuckDbFactCollector.Waits.cs:405 blocking chain facts, by event_time helper analysis facts
8 Lite/Services/LocalDataService.Blocking.cs:593 GetAlertCountsAsync blocking count, by event_time helper alerts, tab badges
9 Lite/Services/LocalDataService.Blocking.cs:599 GetAlertCountsAsync latest event_time helper alerts, tab badges
10 Lite/Services/LocalDataService.Blocking.cs:658 and :659 GetRecentBlockedProcessReportsAsync: the alert engine's form by collection_time (:658), every other caller by event_time (:659) helper Blocking grid, slicers, drill-down, MCP, alert detail rows
11 Lite/Services/LocalDataService.Blocking.cs:863 GetBlockingPairRowsAsync, by event_time helper blocking chain view
12 Lite/Services/LocalDataService.Blocking.cs:911 GetBlockingSlicerDataAsync, by event_time helper blocking slicer
13 Lite/Services/LocalDataService.Blocking.cs:1019 GetBlockingTrendAsync, by event_time helper blocking trend, correlated timeline, MCP
14 Lite/Services/LocalDataService.BlockingStats.cs:98 GetBlockingDurationStatsAsync, by event_time helper blocking severity chart, MCP
15 Lite/Services/LocalDataService.DailySummary.cs:56 daily summary bpr CTE, by event_time helper daily summary, calendar, MCP
16 Lite/Services/LocalDataService.LongQueries.cs:41 GetRecentLongQueryCompletionsAsync, by collection_time helper Long Queries grid
17 Lite/Services/LocalDataService.LongQueries.cs:70 GetSlowestLongQueryCompletionsAsync, by collection_time helper MCP get_long_query_completions
18 Lite/Services/LocalDataService.Overview.cs:117 GetServerSummaryAsync blocking count, by event_time helper Overview cards, MCP
19 Lite/Services/LocalDataService.SystemEvents.cs:329 ReadSystemHealthEventXmlAsync, by event_time and type helper the 8 system_health parsers, in grids and MCP
20 Lite/Services/LocalDataService.SystemEvents.cs:562 CountSystemHealthEventsAsync, by event_time and type helper MCP system_health parser counts
21 Lite/Services/LocalDataService.BlockingStats.cs:57 HasAnyBlockingCaptureAsync, EXISTS with no time window direct the empty-result status of blocking reads
22 Lite/Services/LocalDataService.SystemEvents.cs:510 GetLastSystemHealthCaptureAsync, MAX(collection_time) with no time window direct MCP system_health parser status
23 Lite/Services/LocalDataService.SystemEvents.cs:535 GetLastSystemHealthCaptureOfTypeAsync, the same for one type direct MCP system_health parser status

The three direct reads have no time window, and they read every stored row on purpose:

  • 21: a stored copy cannot change whether a row exists.
  • 22 and 23: a later batch that stored only copies of events already held does move MAX(collection_time), on purpose. That batch still read the session, so its collection_time is a true capture time.

Lite.Tests/StoredEventCopiesSweepTests.cs fails on any other read of the three views, and on a listed read that no longer exists. It also fails when a call puts a collection_time lower bound in its own filter instead of passing it to the helper. No read in LocalDataService.DatabaseStates.cs touches these views. The sources for that file are chosen in #4917.

Performance

All numbers come from Lite's own engine, DuckDB.NET 1.5.5, at Lite's memory_limit of 1 GB. The data is in an in-memory database, with views built the way Lite builds them. The helper's SQL comes from the built Lite assembly. Times are in milliseconds.

The real archive files

These are real rows only: the real archive files, read through read_parquet. The live tables are not read. Each window ends at the busiest server's last archived row, on the read's own clock. Before is the plain view read with the read's own filter, as on dev. After is the same read through the helper. Each figure is the median of 5 runs after a warm-up.

# Read 24 h before 24 h after 30 d before 30 d after
1 Anomaly current_blocking count 6.7 16.8 7.2 41.4
2 Blocking baseline events 7.6 20.1 8.7 36.4
3 Blocking per minute 7.6 15.9 6.3 30.0
4 Top blocking chains 20.0 43.0 17.9 71.6
5 Reconstructed chains 78.9 186.8 97.1 259.3
6 Blocking facts 16.9 40.1 31.6 103.9
7 Blocking chain facts 77.8 175.0 96.9 228.3
8 Alert blocking count 17.6 50.0 13.7 70.7
9 Alert latest event_time 12.7 37.3 12.7 69.1
10 Recent blocked process reports, grid 149.2 283.1 207.5 393.9
10 Recent blocked process reports, alert engine 148.5 350.5 173.7 381.9
11 Blocking pair rows 80.4 172.7 105.3 277.7
12 Blocking slicer 22.1 55.2 76.9 165.0
13 Blocking trend 16.2 48.0 15.5 74.9
14 Blocking duration stats 24.8 61.8 19.2 92.9
15 Daily summary 16.0 40.3 21.3 100.9
18 Server summary blocking count 14.8 33.4 15.9 80.1
19 system_health event XML 30.5 79.8 46.8 107.5
20 system_health event count 25.2 68.2 33.0 103.2

The 24-hour windows held 12 blocked process reports and 1,437 system_health events, with no copies. The 30-day windows held 940 to 942 blocked process reports with 7 copies, or 898 with 7 copies on the alert engine's collection_time window. They held 5,189 system_health events with 9 copies. Reads 21 to 23 run the same SQL as before.

No long query completion archive file exists on this machine, so reads 16 and 17 have no real-file numbers. They use the same helper as the other reads.

A plain COUNT(*) never reads the XML column. The helper has to read and hash the XML in each of its two scans, and a blocked process report's XML averages 13,111 characters here. For the 30-day count of read 1:

Query ms
plain count 10.7
plain count that also reads the XML 38.9
the helper's count 73.5

The cost grows with the rows and the XML inside the read's own window, not with the size of the archive.

100 times the busiest server's month

Real rows plus synthetic rows. The data is the busiest server's 942 real rows from the last 30 days and the look-back. Each has 99 synthetic copies, for 93,258 synthetic rows. Each synthetic copy shifts event_time by a few microseconds and appends its own number to the XML as a comment. So every copy is a distinct event with its own text, while the 7 stored copies inside each set stay copies. That makes 94,200 rows with about 1.2 GB of XML, in one Parquet file.

Each figure is the median of 3 runs after a warm-up, with one form per process and no other test running.

Form 1: anomaly count 2: summary count 3: grid, top 200 4: alert engine, top 200 Slowest run Peak working set
plain read, as on dev 6 19 1,470 1,878 1,491 MB
plain count 7 23 22 21
plain rows: every column of every row, XML included 4,318 4,664 4,842 4,583 4,955
window, XML compared exactly out of memory out of memory out of memory out of memory 1,395 MB
window, md5_number of the XML out of memory out of memory out of memory out of memory 1,025 MB
grouped, md5_number of the XML 15,270 14,292 16,742 15,411 18,045 332 MB
grouped, hash() of the XML (this change) 4,717 4,786 4,629 4,649 5,015 329 MB
  • The plain rows ran in the same process as the plain reads, and its peak includes the rows' .NET strings.
  • Both windows failed on every read with Out of Memory Error: failed to pin block of size 256.0 KiB (953.6 MiB/953.6 MiB used).
  • Both grouped forms returned the same rows: 94,200 to 93,500 for reads 1 to 3, and 89,800 to 89,100 for read 4. The 7 copies in each set went.
  • At this scale, a count through the helper costs about as much as reading every row's XML: 0.96 to 1.09 times plain rows. The rule hashes each row's XML once in each of its two scans. The file is one Parquet row group, so DuckDB reads it on one thread.
  • Why hash(), not md5_number: at 20 times the month, there is 236 MB of XML. One pass over it took 1,718 ms with md5_number, 157 ms with hash() and 111 ms with strlen. DuckDB 1.4.4 took 605 ms for the same md5_number pass.
  • hash() reads the whole text. On DuckDB 1.5.5, 20 texts gave 20 different hashes. One was a base text of 13,000 characters. 18 others each changed one character of it, at places from the first to the last. The last one added a character at the end.

Pins

All of these run on DuckDB.NET 1.5.5.

  • Lite.Tests/StoredEventCopiesTests.cs, for each of the three tables:
    • An event stored again after the reset reads once, and the earliest copy is the one kept.
    • Identical rows from one batch all read, while the same row from a later batch is dropped.
    • Two different events at the same time both read. Their texts have the same length and differ in one character.
    • Two events whose 13,000-character texts differ in one character in the middle both read.
    • Rows with no time, and rows with no text, are never collapsed.
    • A row with NULL text reads next to a twin from a later batch that matches it on every other part. So does a row with empty text.
    • A window that starts between an event's first copy and its later copy reads neither.
    • One more test builds each collector's query with no watermark. It checks that the cutoff is CollectorContext.EventFallbackWindow back, and that the helper's SQL looks back the same number of seconds.
  • Lite.Tests/StoredEventCopiesTests.cs: a long query completion with a NULL database, session or event sequence reads once. Completions that differ only in which of those is NULL stay apart.
  • Lite.Tests/StoredEventCopiesReaderTests.cs: an event's first copy is in an archive file and its later copy is in the live table. Six readers and the alert engine's form of read 10 each count the event once. A window that starts between the two copies shows neither, in both forms of read 10.
  • Lite.Tests/StoredEventCopiesSweepTests.cs: the sweep described above.
  • Lite.Tests/EventBaselineCoveredDaysTests.cs: a blocked process report stored again by a later batch counts once in the blocking baseline.
  • Lite.Tests/WatermarkReadFailureLogTests.cs: each failed watermark read logs one WARN that names the read, the table, the server and the exception type.
  • Darling/Darling.Tests/DarlingEventBaselineCoveredDaysTests.cs and ConsumedTimestampFrameDisciplineTests.cs read the helper's calls in Lite's source.

Red proofs

For each row, the product change was undone, the tests ran and failed, the file was restored, and git status was clean.

At b7dd51435, before the merge with dev:

Product change undone Tests that failed
The long query reads put their collection_time lower bound in their own filter 1: the sweep's lower-bound test
The long query identity has no database, session or event sequence 1: the NULL database, session or sequence test
The join keeps every batch (>= instead of = on the first batch's collection_time) 5 of 5: the earliest copy for each table, the NULL database test and the baseline test
OnEventTime swaps every collection_time in the events CTE, as it did before 1: the baseline test
The blocking baseline reads the plain view 1: the baseline test
The alert engine's form puts its lower bound in its own filter, so the rule has no look-back 2: the reader window test and the sweep's lower-bound test
No XML part in the identity 2 of 3: different events at the same time, for reports and system_health
The XML is keyed by its length 2 of 3: the same test
The XML is keyed by the hash of its first 6,000 characters 2 of 3: the long texts that differ in the middle
The hash also covers the row's id 2 of 3: the earliest copy, for reports and system_health
= instead of IS NOT DISTINCT FROM in the join 3 of 3: the earliest copy
The never-collapse parts test the hash for NULL 3 of 6: the NULL twins
The never-collapse parts skip the empty-text test 3 of 6: the empty-text twins
MAX instead of MIN 3 of 3: the earliest copy
No look-back before the window start 5 of 5: both window-start tests and the shared-window test
No never-collapse parts 12 of 12: no time, no text, and the NULL and empty-text twins

At 2c797f05f, before the dev merge and the change of form:

Product change undone Tests that failed
One reader reads the view directly 3: the sweep and both reader tests
One collector falls back 15 minutes 1: the shared-window test
The helper's look-back is a literal 600 seconds, and the shared window changes to 15 minutes 1: the shared-window test
The first watermark-read catch logs nothing 2 of 3

Tests

  • Lite, 13 classes: the pins above, both sweeps (StoredEventCopiesSweepTests and Lite advice, High Impact Queries and the Agent check read archived rows after the 512 MB archive step #4921's ArchivableTableBareReadSweepTests), and the classes that read these views. 129 tests, 0 failed, at 2a0d4bd49, the merge with dev at 3c5621b47. Neither sweep finds a bare read of the three tables.
  • Darling, DarlingEventBaselineCoveredDaysTests and ConsumedTimestampFrameDisciplineTests: 22 tests, 0 failed, at 2a0d4bd49.
  • Lite full suite at 2a0d4bd49: 7,037 tests, 0 failed. After that head, the Lite source changed only in a doc comment.
  • CI at 2a0d4bd49 failed one Darling test: RepoFileAdoptionTests.TheLfReadingPins_AreExactlyTheOnesDeclaredHere. ConsumedTimestampFrameDisciplineTests now reads StoredEventCopies.cs through the LF reader, and that test's list of LF readers did not name it. e80edd684 adds it to the list.
  • At c6a4dcea7: Darling RepoFileAdoptionTests 2 tests and ConsumedTimestampFrameDisciplineTests 17 tests, 0 failed. Lite StoredEventCopiesTests, StoredEventCopiesSweepTests and StoredEventCopiesReaderTests: 33 tests, 0 failed.
  • Darling full suite at c6a4dcea7: 19,704 tests, 0 failed, 1,265 skipped, 1 not run. The skipped tests need something this run did not have, almost always a live PostgreSQL. CI runs the live PostgreSQL ones.
  • Lite Debug build at c6a4dcea7: 0 warnings.

CHANGELOG

The changelog line is written from this entry at release, so CHANGELOG.md is not edited here.

SECTION: Fixed
ENTRY: Splice into the [#4887] line, in place of "Duplicates stored by earlier resets are not removed.": Blocked process reports, long query completions and system_health events that an earlier reset stored twice now show and count once ([#4918]).
REF: [#4918]: #4918

…ent, long query completion, CPU sample and memory pressure event once after a 512 MB reset re-collects it

The first collection cycle after the reset can read an empty watermark and fetch its fallback window again, so the archive and the hot table both held the same event. The views for these five tables now keep one row per exact identity, the earliest stored copy, and never collapse a row with no usable identity. The FinOps reserved-capacity CPU check now reads the archive view instead of the hot table.
An exact copy of a CPU sample changes no average, maximum or chart line, and a window over that table would cost every read, so cpu_utilization_stats keeps the plain union. Each of the four watermark reads now logs one WARN naming the read, the table, the server and the exception type when it fails and the collector falls back to its fallback window.
… a sweep pins every bare read

FinOps' long-running jobs and file I/O checks, High Impact Queries and the Agent status header read their v_ views, so rows that archival moved to Parquet still count. A source sweep over every archivable table fails on any new bare read that is not on its list of reads that are bare on purpose.
…ad's own filter, and the collectors share their fallback window with that read
…dedupe-after-reset

# Conflicts:
#	Lite/Analysis/AnomalyDetector.cs
#	Lite/Analysis/DrillDownCollector.Blocking.cs
#	Lite/Analysis/DuckDbFactCollector.Waits.cs
#	Lite/Services/LocalDataService.Overview.cs
Dev moved the blocking reads to event_time (#4909, #4913). A read that now
windows on event_time calls StoredEventCopies without collectedFrom and
keeps its event_time bounds in its filter: a stored copy keeps its first
copy's event_time, so both are inside the window or both are outside it.
The alert engine's read of blocked process reports still windows on
collection_time, so it keeps collectedFrom.

OnEventTime now swaps only the events CTE's own local-clock keys and
window. Swapping every collection_time in the CTE would also rewrite the
copy rule inside the blocking baseline's source, and that rule must stay
on collection_time.
The event-baseline twin pin accepts a StoredEventCopies read as Lite's
blocking event source, inside the OnEventTime wrapper, and pins its text.
The consumed-timestamp census counts a StoredEventCopies.<Method>( call
as a read of that method's table. The Lite sweep masks comments and
string bodies with the shared source walker, as CommentFilterAdoptionTests
requires, instead of dropping "//" lines.
A sweep fails when a call puts a collection_time lower bound in its where
instead of passing it as collectedFrom, which would cut the look-back off.
A long query completion with a NULL database, session or event sequence
reads once, and completions that differ only there stay apart. The reader
tests also cover the alert engine's collection_time read of blocked
process reports, and the blocking baseline counts a stored copy once.

The doc says a read on event_time needs no look-back, and states the
look-back's clock assumption: a monitored server clock that is ahead of
the collector's can put a first copy further back than the window.
The rule's key held each row's whole XML. Keying the XML part on md5_number holds 16 bytes per row instead. Every other part stays exact, and so do the CASE parts that keep a NULL or empty text apart. Two different reports would have to collide in 128 bits at one server and event time to merge, and a merge only hides a row from a read.
A window carries every row, its XML included, through the operator. At memory_limit 1GB it ran out of memory on 94,200 blocked process reports. The rule is now each identity's MIN(collection_time), grouped over the identity's keys alone and joined back to the rows with IS NOT DISTINCT FROM on every part. The XML part is DuckDB's 64-bit hash() of the text, named in one place. The parts that keep a row with no text apart test the raw text, because hash(NULL) is not NULL.
The comment said most of their readers filter on collection_time and that a window would run in the view. Most of them filter on event_time now, and the rule is a grouped minimum after each read's own filter.
…among the LF readers

Its StoredEventCopies census now reads that class through the LF reader, on an anchor that crosses the line break between a helper's => and its Read("v_ call.
…oves

The two last-capture reads return MAX(collection_time), which a later batch's copies move on purpose; only the EXISTS read is one a copy cannot change.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 1, 2026 21:50
@erikdarlingdata
erikdarlingdata merged commit 01aa7bf into dev Oct 1, 2026
19 of 20 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/lite-archive-view-dedupe-after-reset branch October 1, 2026 21:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant