Skip to content

Darling reads a wait stored under both spellings as one wait under its clean name - #4941

Merged
erikdarlingdata merged 4 commits into
devfrom
fix/wait-name-trailing-space-history-darling
Oct 2, 2026
Merged

erikdarlingdata merged 4 commits into
devfrom
fix/wait-name-trailing-space-history-darling

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

What changes

From 3.9 the collector trims wait names before it stores them (#4884). SQL Server reports four wait names with a trailing space. A wait that accrued time under both spellings has its history from before the upgrade stored as NAME and its newer rows stored as NAME. Darling's reads keyed on the stored text, so the same wait showed as two rows or two series, each with part of its time. A top-wait pick compared those parts, not the whole wait, and a lookup by the clean name lost the older history.

Now every Darling read that groups, filters or looks up by SQL Server wait name treats both spellings as one wait. It shows that wait under the clean name. There is no schema, migration or continuous-aggregate change.

  • Wide reads that group wait_stats: two levels. The inner query sums per stored name, exactly as before. The outer query adds those sums up per rtrim(wait_type), so the trim runs once per group, not once per row. "Grouping cost" below has the measurements.
  • Trend reads, and reads that group waiting_tasks: rtrim(wait_type) AS wait_type in the select list, and rtrim(wait_type) in GROUP BY and PARTITION BY.
  • Lookups by name: the column stays bare. The read matches wait_type IN ($n, $n || ' '), and a list of several names adds the spaced form of each one.
  • Custom Views: a new catalog flag, TrailingSpaceHistory, marks the three SQL Server wait-name dimensions (wait_stats, waiting_tasks and query_snapshots). The compiler groups them on rtrim(f.wait_type) and keeps the column bare in every filter. It widens the value instead (details below). The pg_wait_stats dimension holds PostgreSQL wait events, so it does not change.

The trailing space

A SELECT-only check of sys.dm_os_wait_stats on SQL Server 2022 and 2025 found these four names. Each one ends in exactly one trailing character, U+0020: EXTERNAL_GOVERNANCE_ATTR_SYNC_BACKGROUND, EDC_DOPP_LOCK, EDC_DOPP_BACKGROUND and SQP_STATS_REPORTING. No other name in that DMV ends in whitespace, so name and name + ' ' cover every spelling the store can hold.

The same check found five names that end in a comma and one name with a space inside it. They were checked and are out of scope. No list in the code holds any of them. The trim does not change them, so the store holds each one with a single spelling.

Scope

Observed:

  • The collector stores only waits with time above zero, and all four sat at 0 ms on both versions in that check.
  • Darling ignores SQP_STATS_REPORTING at collection through IgnoredWaitDefaults.All, so it never stores a clean twin of that one. By design, Darling applies its ignore list at collection only, and its reads have none. So stored history of the spaced form is unchanged by this PR. The history still shows, and it is now labelled with the clean name.

Inferred:

  • Only a wait that accrued time under both spellings is split. That takes EXTERNAL_GOVERNANCE_ATTR_SYNC_BACKGROUND, EDC_DOPP_LOCK or EDC_DOPP_BACKGROUND gaining wait time on a server both before and after its upgrade.

Every read checked

Changed (13 reads, plus Custom Views). "Two levels" means GROUP BY wait_type inside, then rtrim(wait_type) AS wait_type and GROUP BY rtrim(wait_type) over those groups.

Read Where Kind Form
Viewer wait picker ViewerDataService.DistinctWaitTypesSql group Two levels
Viewer wait trends ViewerDataService.WaitTrendsSql filter and group wait_type IN ($4, $4 || ' ', $5, $5 || ' ', ...). rtrim in the select list and the LAG PARTITION BY, so the outer grouping reads the clean name
Viewer current-waits trend ViewerDataService.WaitingTaskTrendSql group rtrim in the select list, GROUP BY and ORDER BY
Viewer FinOps wait categories ViewerDataService.WaitCategorySummarySql group Two levels. The category CASE now reads the clean name, so both spellings always land in one category. None of the four names matches a category, so they land in Other, as before
Viewer "queries with this wait" ViewerDataService.QuerySnapshotsByWaitTypeSql filter wait_type IN ($4, $4 || ' ')
MCP get_wait_stats DarlingDataReader.WaitStatsSql group Two levels
MCP get_wait_types DarlingDataReader.DistinctWaitTypesSql group Two levels
MCP get_wait_trend, both the per-collection and the bucketed form DarlingDataReader.WaitRawCte filter wait_type IN ($2, $2 || ' '). One wait per call, so the LAG needs no partition and runs across the spelling change
MCP get_current_waits_trend DarlingDataReader.WaitingTaskTrendSql group rtrim in the select list, GROUP BY and ORDER BY
Daily summary top wait (viewer calendar, MCP daily-summary tools, fleet sweep) DailySummarySql.RangeSql group Two levels: wait_per_spelling sums per day and stored name, then wait_per_type adds those sums up per day and rtrim(wait_type). The routed forms swap only the queries CTE, so they get it too
Analysis wait facts PgFactCollector.WaitStatsSql group Two levels
Anomaly wait contributors PgAnomalyDetector.WaitContribWindowSql group Two levels
Custom Views on wait_stats, waiting_tasks and query_snapshots ComposeCompiler (GroupRef, BuildFilterClause) group and filter The top-N CTE select list and GROUP BY, the top-N membership test, and the series select list and GROUP BY use rtrim(f.wait_type). Filters keep f.wait_type bare

Custom Views filters on a wait-name dimension:

  • eq and neq bind each value and the value plus one space: f.wait_type = ANY($n) and f.wait_type <> ALL($n).
  • like becomes (f.wait_type LIKE $n OR f.wait_type LIKE $n || ' '), so an exact pattern also matches the spaced spelling.
  • gt becomes (f.wait_type > $n AND f.wait_type <> $n || ' '), and lte becomes (f.wait_type <= $n OR f.wait_type = $n || ' ').
  • gte and lt stay as they were, because a name with one trailing space sorts directly after its clean form. Two guard tests hold this in place (see "Tests"). CI creates its PostgreSQL cluster without --locale (build.yml:1005), and the managed store's initdb passes --locale=C (DarlingManagedPostgres.cs:3440). The order was checked under the C, musl, glibc en_US.UTF-8, ICU and Windows collations.

Checked, no change:

  • Reads that match LCK% only: ViewerDataService.LockWaitTrendSql, DarlingBlockingTrendReader (LockWaitTrendSql, LockWaitTypesSql) and PgDrillDownCollector.LockModeBreakdownSql.
  • Name lists that hold none of the four names:
    • DarlingAlertReadAdapter.PoisonWaitsSql (THREADPOOL, RESOURCE_SEMAPHORE*)
    • the long-running query exclusions (WAITFOR, BACKUP*, XE_LIVE_TARGET_TVF, SP_SERVER_DIAGNOSTICS)
    • the excluded list in PgAnomalyDetector.YoungBaselineBarPeakSql
    • QueryStoreClutter (QDS\_% only)
  • Per-row reads that show each stored row as it is:
    • DarlingSessionReader (ActiveQueriesSql, WaitingTasksSql)
    • PgPileupSnapshotReader
    • PgDrillDownCollector.QueriesAtSpikeSql
    • the viewer's latest-snapshot grids
    • SameStatementPileupDetector
  • Sums over all waits, with no wait key:
    • the overview TotalWaitTrendSql
    • the baseline continuous aggregates (wait_stats_interval_baseline and the older wait_stats_baseline)
    • the signal_wait_pct custom-alert template
  • Significant waits: parsed from the stored system_health XML when read, and the parser trims.
  • Collector stall probes: one top wait per probe row.
  • DarlingDeltaCalculator.WaitStatsSeedSql: Wait names lose their trailing space, and two Hyperscale timer waits are ignored #4884 already accepts the first trimmed sample as a new baseline.
  • dmv_blocking_snapshot: its collector never trimmed, so it holds one spelling before and after.
  • Mute rules: configuration text, not stored wait rows.
  • DarlingFleetReader: it has no wait read. Its only wait_type reference is the mute-rule column test m.wait_type_pattern IS NULL.
  • Web endpoints and MCP tool classes: they dispatch to the reads above.
  • ResolveStoredSnapshotForActualPlanSql: it does not touch wait_type.

Skipped, because these are PostgreSQL wait events and not SQL Server wait names: PgTargetAnomalyDetector, PgTargetBaselineProvider, PgTargetFactCollector.Waits, DarlingPgWaitReader, DarlingPostgresAlertReadAdapter, and the Custom Views pg_wait_stats.wait_type dimension.

Query plan for the lookup

wait_type has no index of its own. The plan check used a fresh PostgreSQL 18.6 database with TimescaleDB 2.30.1:

  • This build's migrations ran on it (156), then the hypertable and compression setup the service runs at start.
  • It held 50 servers x 40 waits x 10 days at 5-minute collections: 5,762,000 rows in 11 daily chunks.
  • The 9 chunks older than a day were compressed, as the compression policy does.
  • EDC_DOPP_LOCK was stored with the space for the first 5 days and clean after that.
  • The table was analyzed before the run.

The query is the get_wait_trend read for one server over an 8-day window, run once with the old wait_type = $2 and once with the new wait_type IN ($2, $2 || ' ').

Both forms take the same path on every chunk:

  • On compressed chunks: an index scan on the compressed chunk's server_id index.
  • On the uncompressed chunks: a bitmap scan of the chunk's (server_id, collection_time) index.
  • wait_type is only ever a filter, and no plan has a sequential scan.

A first run with 5 servers (576,200 rows) took another path on the compressed chunks, the same one for both forms. Each compressed chunk then held about 60 batch rows, so both forms read those small tables sequentially. The uncompressed chunks used the collection_time index. With 50 servers, both forms use the server_id index shown here.

The old form returns 1,441 rows, the clean days only. The new form returns 2,305, both spellings. With plan_cache_mode = force_generic_plan the two forms again match each other. Compressed chunks use the same index, and uncompressed chunks use the chunk's collection_time index, with server_id and wait_type as filters.

Old form, trimmed to the scan nodes:

Append (actual rows=1441.00 loops=1)
  ->  Custom Scan (ColumnarScan) on _hyper_1_3_chunk (actual rows=0.00 loops=1)
        Vectorized Filter: ((collection_time >= '2026-09-24 04:34:00'::timestamp without time zone) AND (collection_time <= '2026-10-02 04:34:00'::timestamp without time zone) AND (wait_type = 'EDC_DOPP_LOCK'::text))
        ->  Index Scan Backward using _hyper_1_3_chunk_compressed_server_id__ts_meta_v2_first_col_idx on _hyper_1_3_chunk_compressed (actual rows=10.00 loops=1)
              Index Cond: ((server_id = 3) AND (_ts_meta_v2_first_collection_time >= '2026-09-24 04:34:00'::timestamp without time zone) AND (_ts_meta_v2_last_collection_time <= '2026-10-02 04:34:00'::timestamp without time zone))
  ->  Custom Scan (ColumnarScan) on _hyper_1_4_chunk (actual rows=0.00 loops=1)
        Vectorized Filter: (wait_type = 'EDC_DOPP_LOCK'::text)
        ->  Index Scan Backward using _hyper_1_4_chunk_compressed_server_id__ts_meta_v2_first_col_idx on _hyper_1_4_chunk_compressed (actual rows=12.00 loops=1)
              Index Cond: (server_id = 3)
  ... chunks 5 to 9: the same shape
  ->  Bitmap Heap Scan on _hyper_1_10_chunk (actual rows=288.00 loops=1)
        Filter: (wait_type = 'EDC_DOPP_LOCK'::text)
        ->  Bitmap Index Scan on _hyper_1_10_chunk_idx_wait_stats_time (actual rows=11520.00 loops=1)
              Index Cond: (server_id = 3)
  ->  Bitmap Heap Scan on _hyper_1_11_chunk (actual rows=55.00 loops=1)
        Filter: (wait_type = 'EDC_DOPP_LOCK'::text)
        ->  Bitmap Index Scan on _hyper_1_11_chunk_idx_wait_stats_time (actual rows=2200.00 loops=1)
              Index Cond: ((server_id = 3) AND (collection_time >= '2026-09-24 04:34:00'::timestamp without time zone) AND (collection_time <= '2026-10-02 04:34:00'::timestamp without time zone))

New form, trimmed to the scan nodes:

Append (actual rows=2305.00 loops=1)
  ->  Custom Scan (ColumnarScan) on _hyper_1_3_chunk (actual rows=234.00 loops=1)
        Vectorized Filter: ((wait_type = ANY ('{EDC_DOPP_LOCK,"EDC_DOPP_LOCK "}'::text[])) AND (collection_time >= '2026-09-24 04:34:00'::timestamp without time zone) AND (collection_time <= '2026-10-02 04:34:00'::timestamp without time zone))
        ->  Index Scan Backward using _hyper_1_3_chunk_compressed_server_id__ts_meta_v2_first_col_idx on _hyper_1_3_chunk_compressed (actual rows=10.00 loops=1)
              Index Cond: ((server_id = 3) AND (_ts_meta_v2_first_collection_time >= '2026-09-24 04:34:00'::timestamp without time zone) AND (_ts_meta_v2_last_collection_time <= '2026-10-02 04:34:00'::timestamp without time zone))
  ->  Custom Scan (ColumnarScan) on _hyper_1_4_chunk (actual rows=288.00 loops=1)
        Vectorized Filter: (wait_type = ANY ('{EDC_DOPP_LOCK,"EDC_DOPP_LOCK "}'::text[]))
        ->  Index Scan Backward using _hyper_1_4_chunk_compressed_server_id__ts_meta_v2_first_col_idx on _hyper_1_4_chunk_compressed (actual rows=12.00 loops=1)
              Index Cond: (server_id = 3)
  ... chunks 5 to 9: the same shape
  ->  Bitmap Heap Scan on _hyper_1_10_chunk (actual rows=288.00 loops=1)
        Filter: (wait_type = ANY ('{EDC_DOPP_LOCK,"EDC_DOPP_LOCK "}'::text[]))
        ->  Bitmap Index Scan on _hyper_1_10_chunk_idx_wait_stats_time (actual rows=11520.00 loops=1)
              Index Cond: (server_id = 3)
  ->  Bitmap Heap Scan on _hyper_1_11_chunk (actual rows=55.00 loops=1)
        Filter: (wait_type = ANY ('{EDC_DOPP_LOCK,"EDC_DOPP_LOCK "}'::text[]))
        ->  Bitmap Index Scan on _hyper_1_11_chunk_idx_wait_stats_time (actual rows=2200.00 loops=1)
              Index Cond: ((server_id = 3) AND (collection_time >= '2026-09-24 04:34:00'::timestamp without time zone) AND (collection_time <= '2026-10-02 04:34:00'::timestamp without time zone))

The viewer drill-down lookup uses the same IN form on query_snapshots, whose indexes have the same shape: (server_id, collection_time) and (collection_time DESC).

Grouping cost

rtrim in GROUP BY runs once per row. The two-level form keeps the old per-chunk aggregation and trims once per group. Each form ran on the same 5,762,000-row database, for one server, with EXPLAIN (ANALYZE, BUFFERS) and plan_cache_mode = force_custom_plan. After one warm-up, each form ran 3 times, with the forms interleaved. The table shows the median:

Read Window Before this PR, ms One level, ms Two levels, ms
DarlingDataReader.WaitStatsSql 7 days 27.1 35.2 (+30%) 28.6 (+5%)
DailySummarySql.RangeSql 10 days 38.8 47.5 (+22%) 39.5 (+2%)

That first set timed a two-level WaitStatsSql with other inner column names. A second set used the exact shipped text and kept the same order. It gave 32.5, 39.9 and 31.3 ms for WaitStatsSql, and 45.1, 57.0 and 41.9 ms for the range read.

The one-level form was more than 10% slower in both reads. So every wide grouping read over wait_stats now uses two levels:

  • both DistinctWaitTypesSql reads
  • DarlingDataReader.WaitStatsSql and PgFactCollector.WaitStatsSql
  • PgAnomalyDetector.WaitContribWindowSql
  • the range read in DailySummarySql
  • ViewerDataService.WaitCategorySummarySql

Each outer aggregate is a sum of the inner sums, so it adds up to the same total. The one-level and two-level forms returned the same rows for all 50 servers. That is 2,000 rows for WaitStatsSql and 500 for the range read, with no difference either way (EXCEPT ALL on the row text).

No form has a VectorAgg node on this TimescaleDB 2.30.1 store. All three keep the per-chunk partial aggregation. Plan nodes from the second set, trimmed, with chunk numbers shown as N:

WaitStatsSql, before this PR
Finalize HashAggregate  Group Key: wait_stats.wait_type
  8x Partial HashAggregate  Group Key: _hyper_N_chunk.wait_type
    6x Custom Scan (ColumnarScan) <- Index Scan using _hyper_N_chunk_compressed_server_id__ts_meta_v2_first_col_idx
    2x Bitmap Heap Scan <- Bitmap Index Scan on _hyper_N_chunk_idx_wait_stats_time

WaitStatsSql, one level
Finalize HashAggregate  Group Key: (rtrim(wait_stats.wait_type))
  8x Partial HashAggregate  Group Key: rtrim(_hyper_N_chunk.wait_type)
    (the same scans)

WaitStatsSql, two levels
GroupAggregate  Group Key: (rtrim(per_spelling.wait_type))
  Subquery Scan on per_spelling
    Finalize HashAggregate  Group Key: wait_stats.wait_type
      8x Partial HashAggregate  Group Key: _hyper_N_chunk.wait_type
        (the same scans)

RangeSql, before this PR
Finalize HashAggregate  Group Key: (date_trunc('day', wait_stats.collection_time)), wait_stats.wait_type
  10x Partial HashAggregate  Group Key: date_trunc('day', _hyper_N_chunk.collection_time), _hyper_N_chunk.wait_type
    8x Custom Scan (ColumnarScan) <- Index Scan using _hyper_N_chunk_compressed_server_id__ts_meta_v2_first_col_idx
    2x Bitmap Heap Scan <- Bitmap Index Scan on _hyper_N_chunk_idx_wait_stats_time

RangeSql, one level
Finalize HashAggregate  Group Key: (date_trunc('day', wait_stats.collection_time)), (rtrim(wait_stats.wait_type))
  10x Partial HashAggregate  Group Key: date_trunc('day', _hyper_N_chunk.collection_time), rtrim(_hyper_N_chunk.wait_type)
    (the same scans)

RangeSql, two levels
HashAggregate  Group Key: wait_per_spelling.d, rtrim(wait_per_spelling.wait_type)
  Subquery Scan on wait_per_spelling
    Finalize HashAggregate  Group Key: (date_trunc('day', wait_stats.collection_time)), wait_stats.wait_type
      10x Partial HashAggregate  Group Key: date_trunc('day', _hyper_N_chunk.collection_time), _hyper_N_chunk.wait_type
        (the same scans)

Still one level:

  • ViewerDataService.WaitTrendsSql filters on the chosen names first, so it trims only their rows.
  • The two WaitingTaskTrendSql reads group waiting_tasks, which holds only the tasks waiting at each collection, not a row for every wait type.
  • Custom Views. A panel can aggregate with a percentile, and a percentile cannot be rebuilt from per-group results.

Tests

New WaitNameSpacedHistoryLivePostgresTests: 21 tests, all _AgainstDevPostgres, so CI's PostgreSQL jobs run them. They cover three things:

  • Each changed grouping read returns a spaced row and a clean row as one row with the summed values.
  • Each lookup by the clean name finds the spaced history.
  • Custom Views: grouping, every filter operator, and a top-N time series with the "(other)" series.

They failed on the code without the fix, 16 of 16, each on its assertion:

AnalysisWaitFacts_ANameStoredBothWays_IsOneRowWithTheSummedValues: Expected "EDC_DOPP_LOCK|5|600|50" / Actual "PREEMPTIVE_OS_WRITEFILE|5|500|50"
DailySummary_TopWait_SumsBothSpellingsUnderTheCleanName: Expected "EDC_DOPP_LOCK" / Actual "PREEMPTIVE_OS_WRITEFILE"
McpWaitTrend_ByTheCleanName_IncludesTheSpacedHistory: Expected [(07:53:10, 1), (07:58:10, 2)] / Actual [(07:58:10, 2)]
ViewerCurrentWaitsTrend_ANameStoredBothWays_IsOneRowWithTheSummedDuration: Expected [("EDC_DOPP_LOCK", 1000), ("EDC_DOPP_LOCK", 50)] / Actual [("EDC_DOPP_LOCK", 600), ("EDC_DOPP_LOCK ", 400), ("EDC_DOPP_LOCK ", 50)]
ViewerQueryDrillDown_ByTheCleanName_FindsASnapshotStoredWithTheSpace: Expected [72, 71] / Actual [72]
CustomViewGroupedByWaitType_ANameStoredBothWays_IsOneRowWithTheSummedValue: Expected "EDC_DOPP_LOCK|600" / Actual "PREEMPTIVE_OS_WRITEFILE|500"
McpCurrentWaitsTrend_ANameStoredBothWays_IsOneRowWithTheSummedDuration: Expected [(07:58:32, "EDC_DOPP_LOCK", 1000), (08:01:32, "EDC_DOPP_LOCK", 50)] / Actual [(07:58:32, "EDC_DOPP_LOCK", 600), (07:58:32, "EDC_DOPP_LOCK ", 400), (08:01:32, "EDC_DOPP_LOCK ", 50)]
AnomalyContributors_ANameStoredBothWays_IsOneRowWithTheSummedValue: Expected "EDC_DOPP_LOCK|600" / Actual "PREEMPTIVE_OS_WRITEFILE|500"
McpWaitTypes_ANameStoredBothWays_IsOneCleanNameRankedByTheSum: Expected "EDC_DOPP_LOCK" / Actual "PREEMPTIVE_OS_WRITEFILE"
CustomViewFilteredByTheCleanName_MatchesBothSpellings(op: "eq"): Expected "600" / Actual "300"
CustomViewFilteredByTheCleanName_MatchesBothSpellings(op: "like"): Expected "600" / Actual "300"
CustomViewFilteredByTheCleanName_MatchesBothSpellings(op: "neq"): Expected "500" / Actual "800"
McpWaitStats_ANameStoredBothWays_IsOneRowWithTheSummedValues: Expected [("EDC_DOPP_LOCK", 5, 600, 50), ("PREEMPTIVE_OS_WRITEFILE", 5, 500, 50)] / Actual [("PREEMPTIVE_OS_WRITEFILE", 5, 500, 50), ("EDC_DOPP_LOCK", 2, 300, 20), ("EDC_DOPP_LOCK ", 3, 300, 30)]
ViewerWaitTrends_ByTheCleanName_AreOneSeriesThatIncludesTheSpacedHistory: Expected [07:53:00, 07:58:00] / Actual [07:58:10]
ViewerWaitPicker_ANameStoredBothWays_IsOneCleanNameRankedByTheSum: Expected "EDC_DOPP_LOCK" / Actual "PREEMPTIVE_OS_WRITEFILE"
ViewerWaitCategorySummary_TopWait_SumsBothSpellingsUnderTheCleanName: Expected "EDC_DOPP_LOCK" / Actual "PREEMPTIVE_OS_WRITEFILE"
Total: 16, Failed: 16

The gt and lte cases came later, with the range-operator handling. They failed when the flag was turned off on the wait_stats dimension, which compiles the old filter SQL:

CustomViewFilteredByTheCleanName_MatchesBothSpellings(op: "gt"): Expected "500" / Actual "800"
CustomViewFilteredByTheCleanName_MatchesBothSpellings(op: "lte"): Expected "600" / Actual "300"

The top-N time series test came last. A spaced and a clean row of one wait sit in the top N. With the flag off, the read ranks one spelling at a time. Each spelling (300 ms) loses the one slot to the rival (500 ms), and the whole wait folds into "(other)". It failed that way:

CustomViewTopNSeriesWithOther_ANameStoredBothWays_IsOneSeriesInTheTopN: Expected "(other)=500" / Actual "(other)=600"

The gte and lt cases are guards, not red-first tests. Those two operators compile the same with the flag on or off, and both cases pass either way:

  • gte on the clean name keeps its spaced history.
  • lt on the clean name drops both spellings of that name, and keeps a lower name, EDC_DOPP_BACKGROUND, in both spellings.

With the fix, all 21 pass against a migrated PostgreSQL 18.6 database with TimescaleDB 2.30.1.

Also:

  • New SQL-shape tests in DarlingComposeTests check four things:
    • every filter operator keeps the column bare
    • grouping uses rtrim
    • eq and neq bind both spellings
    • pg_wait_stats has no rtrim, and the catalog flags exactly the three SQL Server wait-name dimensions
  • Text tests that held the old SQL are updated: DarlingMcpDataToolsTests, ViewerDrillDownTests, ViewerWaitStatsTests and the two top-N series tests in DarlingComposeTests.
  • Full Darling.Tests run without a PostgreSQL connection:
    • 19,868 total, 0 failed
    • 1,320 skipped (the live PostgreSQL tests, these 21 among them)
    • 1 not run, an existing explicit test
  • The build has 0 warnings.

CHANGELOG

No new entry: this amends the [3.9.0] #4884 line, and the release owner applies it.

Proposed sentence for Darling:

Darling reads a wait stored with or without the trailing space as one wait under its clean name, so older history stays with it.

…s clean name

SQL Server reports four wait names with a trailing space, and from 3.9 the collector
stores them trimmed (#4884). History from before the upgrade keeps the spaced name,
so the same wait could be stored under two spellings.

Reads that group by wait now key on rtrim(wait_type): the viewer's wait picker, wait
trends, current-waits trend and FinOps category summary, the MCP wait stats, wait
types and current-waits trend, the daily summary's top wait, the analysis wait
facts and the anomaly wait contributors. Lookups by name keep the column bare and
match both spellings: the MCP wait trend and the viewer's queries-by-wait drill-down
use wait_type IN ($n, $n || ' '). Custom Views group the three SQL Server wait-name
dimensions on the trimmed name and widen each filter value instead of wrapping the
column. pg_wait_stats (PostgreSQL wait events) is unchanged.
…ellings on rtrim

rtrim in GROUP BY ran once per row and cost 22-30% on the 7-day wait stats
and 10-day daily summary reads. The six wide grouping reads now keep the old
per-name aggregation inside and merge the two spellings over those groups,
which measures within a few percent of the read before #4941.

The FinOps wait categories read the category from the clean name.

Tests: gte/lt guard cases and a top-N time series with "(other)" in the
live spaced-history class; the SQL text tests follow the two-level shape.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 2, 2026 09:46
@erikdarlingdata
erikdarlingdata merged commit 813bde2 into dev Oct 2, 2026
16 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/wait-name-trailing-space-history-darling branch October 2, 2026 09:46
erikdarlingdata added a commit that referenced this pull request Oct 2, 2026
The 3.9.0 entry for the wait-name fix (#4884) now covers #4939 and #4941, which completed it after the release was cut: its references gain both PRs in the index and the archive, rows stored with the trailing space before the upgrade are described as reading as the clean name in both apps, and the entry says Lite's wait views hide the two newly ignored waits in history from before the upgrade while Darling shows that history until retention removes it. The archive census carries the new 3.9.0 prose hash; the entry count is unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant