Skip to content

Release 3.9.0 - #4937

Merged
erikdarlingdata merged 930 commits into
mainfrom
dev
Oct 2, 2026
Merged

erikdarlingdata merged 930 commits into
mainfrom
dev

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Promotes dev to main for the 3.9.0 release. Full entry: docs/changelog/3.9.md (indexed in CHANGELOG.md), dated 2026-10-02.

Before upgrading, note the release's Important section:

The head is dev. This stays a draft until the release candidate built from this commit has been installed and checked.

erikdarlingdata and others added 30 commits September 28, 2026 00:11
…e waits get a real benefit (#4555)

Plan analysis turns a plan's wait statistics into findings, and gives external and preemptive waits a real benefit. Fixes #4516, fixes #4517, part of #4511.

- BenefitScorer.ScoreWaitStats, matching PerformanceStudio, adds a "Wait: <type>" finding for each significant wait type in an actual plan, with a severity tier from its estimated benefit and a description from WaitStats.json. The file ships as an embedded resource of the shared plan-analysis assembly, so Darling, Lite and the viewer read the same copy.
- External and preemptive waits get PerformanceStudio's separate benefit formula instead of being folded into operator time.
- GetOperatorMaxThreadOwnCpuMs matches PerformanceStudio: it skips thread 0 and looks through batch mode zones and Compute Scalar pass-throughs.
- The findings appear now that the scorer runs at every entry point (#4552). They reach only the plan warning fact and the advice headline's critical count, not any alert or health band.
- Tests: PlanSync4516Tests, PlanSync4517Tests and PlanSync4517ProbeTests, failing without the change; two mutations fail them.
Plan analysis reports which operator each finding came from. Fixes #4534, part of #4511.

- PlanWarning gains OriginNodeIds, ported from PerformanceStudio (9c3fceb). The per-node pass stamps each finding with the node it came from once, at its end, after rule 35, so the Expensive Operator finding carries its node too. The table-variable rule keeps its lists of referencing and modifying nodes.
- The MCP plan analysis output carries origin_node_ids next to source and max_benefit_percent, keeping the benefit ordering.
- Tests: PlanSync4534Tests, and PlanSync4534OriginFixturesTests, which check the IDs against the operators in three of PerformanceStudio's own test plans (added under Darling.Tests/Fixtures/OriginPlans); two mutations fail them.
…k changes (#4553)

Store: the rollup coverage probe re-measures a rollup's floor only when its oldest chunk changes, or after an hour. Fixes #4539.

The probe re-ran min(bucket) on every rollup's materialization every five minutes. On a compressed chunk that sort reads every compressed batch first; on one production store that was about 170 MB of temp every five minutes.

- A per-data-source cache keyed on each rollup's oldest materialization chunk name and materialization hypertable. A cheap catalog read (RollupOldestChunkSql) decides which rollups changed; only those get a fresh min(bucket).
- MergeRollupFloors, a pure function: a rollup measured this cycle always uses the measured value, and a NULL result (empty, or re-created with no data) drops its cache entry. Only rollups not measured this cycle reuse a cached floor, and never for more than an hour (RollupFloorMaxReuse), which bounds how long an out-of-band delete in the oldest chunk could go unseen. A rollup that no longer exists drops its entry.
- A failed catalog read falls back to measuring every rollup, as before. RollupCoverageProbeSql's one-argument overload, used by the retention arming check, is unchanged.
- Tests: RollupFloorCacheTests (the planner, the merge, the hypertable key, the one-hour limit with an injectable clock) and RollupFloorCacheLiveTests (the floor matches a fresh min, a second call runs no min(bucket), a dropped chunk re-measures); IntervalHonestHourlyRollupLiveTests covers a rollup dropped and re-created empty.
…s at any hour (#4558)

Lite tests: the FinOps tests' reads cover their seeded rows at any hour.

TestDataSeeder anchors its period end to the most recent UTC midnight plus 4 hours, so the CPU-spike scenario lines up with a UTC day (#4385), and the high-impact seed stamps its rows 30 minutes before that. FinOpsTests read those rows with a fixed 24-hour lookback from now, so between 03:30 and 04:00 UTC the rows fell outside the window and HighImpactSkew_DominantQueryScoresHighest and HighImpactSkew_DominantQueryHighScore failed.

- RunHighImpactAsync now computes its lookback from now back to TestDataSeeder.TestPeriodStart, plus an hour, unless a test passes one explicitly. The seeder's anchor is unchanged.
- HighImpactWindowCoversSeedAtWorstHour checks the window still covers the seed at 03:59 UTC.
- CleanServer_NoDuckDbRecommendations failed the same way: the utilization read uses a fixed 24-hour window with no parameter, so with no CPU samples in it the engine saw 0% CPU and recommended downsizing. SeedCleanFinOpsServerAsync now also seeds the same flat 50% CPU signal over the last three hours; it stays flat, like the original seed, so the reserved-capacity stability rule (which needs a nonzero standard deviation) does not fire. The other FinOps and recommendation scenarios read 7-day windows and are not affected.
…4559)

Plan viewer: a finding's header links to the operator it came from. Refs #4534, part of #4511.

- The plan-level and operator-level warning headers use each finding's OriginNodeIds, as PerformanceStudio's AttachOriginNavigation and TryNavigateToNode do: a header with origin operators ends in an arrow, shows a hand cursor and a tooltip (one operator or several), and a click selects the operator and scrolls it into view. A finding with no origin operators keeps a plain header.
- The Darling viewer and Lite share PlanViewerControl, so both get it. No new panel or control.
- PlanWarningDisplay gains the header-text helper. Tests: PlanViewerOriginNavigationTests covers none, one and several origin operators; a mutation fails them.
…t parses it (#4551)

Plan analysis: a deeply nested plan no longer crashes the process that parses it. Refs #4512, part of #4511.

ShowPlanParser's recursive walk costs about 12.8 KB of stack per nested operator, so a legitimately deep plan overflowed at about 80 levels on a 1 MB thread (the plan viewer) and about 120 on a 1.5 MB thread (the Darling analysis pass and MCP). A stack overflow cannot be caught, so the process ended, and an analysis pass that met the same stored plan after every restart would crash-loop.

- From PerformanceStudio (e8e5a21, f2153e2): MaxParseDepth (1,000), carried across procedure and UDF sub-plans; MaxParseCharacters (16 MB), checked before parsing; and a catch around the walk that sets ParsedPlan.ParseError.
- Parse runs its walk on a dedicated thread with a 32 MB stack and joins it, so the 1,000-level guard is reachable on any caller. Every failure inside the walk is caught on that thread and returned as ParseError. Analyze, Score and Layout stay on the caller's thread: at depth 999 they need about 64 to 128 KB.
- ParseError is read everywhere: the MCP plan tools report parse_error, the plan-advisory aggregator skips the plan, PlanAnalysisPipeline.Run leaves a partially parsed plan unanalyzed and unscored, both drill-down collectors stop after parsing, and the plan viewer shows "This plan couldn't be parsed: <reason>". Malformed XML sets ParseError too.
- ScopedDescendants is iterative with an explicit stack; the recursive version could overflow on a deep run of non-operator elements and was quadratic.
- Tests: PlanSync4512Tests (depth, sub-plan depth, size, from a real 1 MB thread), PlanSync4512PostParseTests (depth 999 and 150 through parse, analyze, score and layout on 1 MB and 1.5 MB threads; 1,001 refused), PlanSync4512ParseErrorSurfacingTests (MCP, the aggregator, the pipeline and a partially parsed plan, malformed XML, the drill-down guards), and 100,000 nested non-operator elements on a 1 MB thread. The crash on the old code was shown in a separate process, locally.
…plan now get findings (#4557)

Plan analysis finds statements inside a procedure, function or cursor sub-plan. Fixes #4514, part of #4511.

- PlanStatements.EnumerateAll, ported from PerformanceStudio (74ad0e4), walks every statement including those nested in sub-plans, with an explicit stack. The analyzer, the scorer, ComputeOperatorCosts, AllWarnings and AllMissingIndexes, the MCP plan output, both drill-downs and the plan viewer's statement list use it.
- The parser reads sub-plans before its no-plan early return, so an EXEC statement's procedure and function bodies are parsed, and cursor statements read theirs too (PerformanceStudio 4d978b4).
- Stored plan advisories in Darling are unaffected, because cached plans do not nest sub-plans. Plans opened in the viewer, pasted, or sent to analyze_plan_xml gain findings for their nested statements.
- Tests: six cases, failing on the code before the port; a mutation fails them.
…e a function or conversion is on (#4554)

Plan analysis: the non-SARGable check reads which side of a comparison, and which table, a function or conversion is on. Fixes #4518, fixes #4519, part of #4511.

- A CONVERT_IMPLICIT on the parameter side of a comparison is no longer flagged; only a conversion wrapping the column side makes the predicate non-SARGable.
- The side check is scoped to the function's own comparison, and ISNULL and COALESCE are scoped the same way. LIKE counts as a comparison operator.
- ScanIdentity decides which table a column belongs to: a table variable's bare column counts, a Nested Loops outer reference does not, and temp table names are cleaned before matching.
- Ported from PerformanceStudio (1c17b06, 6314f2c, 1016100, 9a04629, 1374b75).
- Tests: PlanSync4518Tests, PlanSync4518EndToEndTests and PlanSync4519Tests; the end-to-end cases fail without the change, and two mutations fail them.
Plan analysis: parsing and analysis stop when the caller cancels. Fixes #4512, part of #4511.

- ShowPlanParser.Parse and a new ParseAsync, PlanAnalyzer.Analyze, BenefitScorer.Score and PlanAnalysisPipeline.Run and RunAsync take a CancellationToken and check it while parsing, per statement and down the node walks, and between pipeline stages (PerformanceStudio 4ab01f1, dc03dc1, 91f8617). A cancellation reaches the caller as the original OperationCanceledException, never as a ParseError.
- ParseAsync runs the same dedicated 32 MB-stack thread parse as Parse, and both share one up-front size check. The parse thread also turns any other exception into a ParseError instead of ending the process.
- The Darling and Lite MCP plan tools, the plan-advisory aggregator (new ExtractCancellable and SummarizeCancellable; Extract and Summarize keep their signatures), both fact collectors and both drill-downs pass their token. The drill-downs' two catches no longer swallow a shutdown cancellation.
- Tests: PlanSync4512CancellationTests (pre-cancelled, cancelled from inside the walk through an internal per-statement hook, the pipeline, the size refusal, the token-free overloads, the abandon classifier) and PlanSync4512AsyncParseTests (a 500-level plan through ParseAsync on a 1 MB thread; the old code's stack overflow was shown in a separate process, locally).
Stop swallowing a cancelled request in 8 catch blocks.

- Darling's DarlingAnalysisService.CollectAndScoreFactsAsync and ComparePeriodsAsync exclude OperationCanceledException from their fault catch, as CollectConfigAuditFactsAsync in the same file already does. A cancelled get_analysis_facts or compare_analysis call now ends as a cancellation instead of an empty result logged as an error.
- Lite's RemoteCollectorService watermark and state reads (GetLastCollectedTimeAsync, GetLastCollectedTimeWithFrameAsync, GetLastCollectedTimeForDatabaseAsync, GetLastCollectedInstanceIdAsync, HasPriorCollectorSuccessAsync) and Darling's DarlingCollectorRunner.HasPriorCollectorSuccessAsync let a cancellation propagate instead of returning a fallback value, such as "no prior success", to a cycle that is shutting down.
- Every fixed read is called only with its request's or collection cycle's own token, so a cancellation there always means the caller gave up. The remaining unfiltered catches that were read and kept are listed in the PR.
- Tests: AnalysisReadCancellationPinTests (a pre-cancelled token through both public Darling methods; reverting either filter fails it) and RemoteCollectorServiceCancellationCensusTests (all six reads; fails on the code before the change).
#4561)

Lite: a cancelled analysis request stops instead of reading to the end.

- Lite's AnalysisService CollectAndScoreFactsAsync, CollectConfigAuditFactsAsync, ComparePeriodsAsync and the private LookUpDispersionAsync take a CancellationToken and set it on AnalysisContext (or pass it to BaselineProvider), which the DuckDB fact collector and baseline provider already observe on every store read. This matches Darling's twin service.
- Each method's fault catch excludes OperationCanceledException, so a cancelled read reaches the caller as a cancellation, not an empty result logged as an error.
- The get_analysis_facts, compare_analysis and audit_config MCP tools take and pass the request's token, and their outer catches exclude a cancellation. AttributeConfigChangesAsync passes its token to its ComparePeriodsAsync call.
- Tests: LiteAnalysisServiceCancellationSourceTests (Darling.Tests) checks each method's token parameter, its filtered catch and that every caller passes a token; a mutation fails it. LiteAnalysisCancellationTests (Lite.Tests) checks that a cancelled token makes the three public methods throw.
…d to match (#4563)

Dependencies: Velopack 1.2.0 to 1.2.158, with the release packer moved to match.

- Directory.Packages.props moves Velopack, the installer and self-updater for Lite and the Darling Viewer, to 1.2.158.
- Every lock file the bump reaches is regenerated with dotnet restore --force-evaluate, so locked-mode restore passes (Lite, Lite.Tests, the Darling Viewer, Darling.Tests, and the Dashboard and Dashboard.Tests lock files).
- The vpk tool that packs each release is installed at 1.2.158 in both release workflows, matching the library.
…ts CPU:Elapsed ratio (#4565)

Plan analysis: drop three rules PerformanceStudio removed, and show its CPU:Elapsed ratio. Fixes #4537, part of #4511.

- Rules 21 (CTE multiple references), 25 and 31 (the parallelism-efficiency warnings) are removed with their helpers and the BenefitScorer cases, as in PerformanceStudio aabbaa2 and 28d4c74. An actual plan's runtime stats show where the time went, so these text- and ratio-based warnings added noise.
- The plan viewer's runtime summary gains PerformanceStudio's CPU:Elapsed row after Elapsed, with external-wait time taken out of CPU first (PlanDisplayText.CpuElapsedRatio). The Darling viewer and Lite share the control.
- The wait-stat half of aabbaa2 already matches: descriptions come from the curated WaitStats.json. The legacy-rule marker from the same commit is its own change, #4566.
- Tests: PlanSync4537Tests (the three removals, which fail on the code before the change, and the ratio including the external-wait subtraction); mutations restoring rules 25/31 or removing the subtraction fail them.
…longer fails the test (#4581)

Tests: a Windows file lock during managed-PostgreSQL test cleanup no longer fails the test.

- DarlingManagedPostgresTests.Postgres17Store_WithTheOldValueInItsManagedConf_StartsWithTheCap_Gated failed twice on Windows CI after its assertions passed: deleting its temp directory threw UnauthorizedAccessException on a file the just-stopped PostgreSQL process still mapped.
- The best-effort temp-directory deletes after a PostgreSQL stop (8 sites in six test files) now catch IOException or UnauthorizedAccessException, as their comments already intended. Test code only; no assertion changes.
- The gated tests skip on this change (the previous-major runtime fixture is built only for store-path changes); the next run that builds it exercises the new catch.
Wait-category colours: a laid-out, contrast-checked palette. Fixes #4578, part of #4511.

- ChartPalette.WaitColor's dark-background wait-category colours match PerformanceStudio dev's palette (DarkTheme.axaml, WaitCategory.*) for all 23 shared categories. Three colours that fell short of 3:1 contrast against the dark chart background are fixed.
- Batch Mode, a PerformanceMonitor-only category, and the light-background table stay PerformanceMonitor's own; the light table was re-checked against the new values.
- The palette is shared code, so the plan viewer and Lite's wait charts both change.
- Tests (Viewer4578Tests): every category matches PerformanceStudio's value, and every colour clears 3:1 against the dark background; both fail on the code before the change.
Plan analysis: the repro script matches PerformanceStudio's. Fixes #4564, part of #4511.

- ReproScriptBuilder ports PerformanceStudio dev's hardening (719d861, 04ae735, 32be789, 9488bf0's declaration-list scan, 37d2ec3, 38ea2b6): a parameter's name and data type from the plan's ParameterList must be a plain identifier and a structural type (one to three parts, an optional size, precision or max) before they are declared, or the parameter is dropped with a header warning; a compiled value that is not a safe literal gets a placeholder; and every header-comment value, the product name included, goes through CommentSafe.
- Beyond PerformanceStudio: parameter names may start with a digit after the @, so the @0, @1 names from simple and forced parameterization keep working; the USE line is left out, with a warning, when the database name contains a control character; and every validation pattern is anchored with \A and \z. PerformanceStudio's copy has the same gaps (erikdarlingdata/PerformanceStudio#590).
- A malformed type is refused with a clear message; SQL Server itself rejects a malformed sp_executesql parameter definition, so this is hardening rather than an injection fix.
- Tests: PlanSync4564ReproScriptTests (the ported cases plus the @1/@2 parameterized plan, a malformed type, control characters in the database name, a trailing newline and a hostile product name), with mutations; ReproScriptBuilderHardeningTests still pass.
…lours from one helper (#4584)

Plan viewer: show each warning's actionable fix, and take severity colours from one helper. Fixes #4572, part of #4511.

- The plan-level and per-operator warning lists render a warning's ActionableFix as an italic line under its message, as PerformanceStudio 3d717e7 does. No rule sets ActionableFix today, in either codebase, so nothing new shows yet; the line is in place for when one does.
- PlanWarningDisplay.WarningSeverityColorHex replaces five copies of the same severity-colour ternary in the properties panel and tooltips, as PerformanceStudio 51d59ca does, with the same colour values.
- The control is shared, so the Darling viewer and Lite both get it.
- Tests (PlanViewerWarningColorAndFixTests): the three severity colours through the helper, and no inline severity colour literal left in the viewer control.
Plan viewer: colour edges by the actual/estimated row ratio. Fixes #4579, part of #4511.

- PlanEdgeColour.ForChild ports PerformanceStudio 70d28e6's link colouring: on an actual plan, an edge whose operator returned more rows than estimated turns light orange, orange or red, and one that returned fewer turns blue, light blue or bright blue, in tiers at the divergence limit, 10 times it and 100 times it; within the limit, and on estimated plans, the edge keeps its plain colour. A zero estimate with actual rows counts as the widest underestimate.
- The divergence limit defaults to 10, as in PerformanceStudio, and is floored at 2. It is a settings-file value in both apps (Lite's settings.json, the Darling Viewer's settings); like PerformanceStudio, there is no Settings-window control for it.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (Viewer4579Tests): every tier boundary, estimated plans, and a zero estimate; they fail on the code before the change.
…, not a fixed group count (#4568)

get_ag_health fills its fleet-wide answer to the 32 KB budget by size, not a fixed group count. Refs #4471, #4474.

- DarlingAgReader.Build walks the availability groups most-severe-first, serializing each group once and adding its bytes to the envelope's, and stops before the group that would cross McpResponseBudget.DefaultBytes. At least one group always comes back.
- The response actually returned is then re-measured, note included, and the last group dropped (the note rebuilt) while it is over the budget.
- The default of 11 groups stays as an upper bound, an explicit limit is still honored, and groups_returned, groups_total, groups_truncated and the note stay truthful; the note says whether the size budget or the limit cut the page.
- The tool and parameter descriptions say so, within the tools-list budget.
- Tests (DarlingAgReaderTests): 42 groups of 10 databases each come back within 32 KB and truncated (280,142 bytes on the code before the change); 5 small groups all come back; one oversized group returns alone, not truncated; an explicit small limit is exact; and groups that fit only without the note are cut to fit with it. The live test checks that the serialized response fits the budget.
)

Plan viewer: the runtime summary card matches PerformanceStudio's: title, order, memory-grant colours, spill flag. Fixes #4570, part of #4511.

- The runtime summary card is titled "Predicted Runtime" for an estimated-only plan, and its rows follow PerformanceStudio dev's order: Elapsed, CPU:Elapsed, DOP or the serial reason, CPU, UDF CPU and elapsed, Compile, the memory grant, Optimization, the CE model and the rest. Compile time always shows.
- The memory-grant efficiency is coloured as in PerformanceStudio (40 percent or more, 20 or more, below 20; over 100 percent is red), and a spill anywhere in the plan caps it at orange and adds a spill flag.
- PlanDisplayText.BuildRuntimeSummaryRows builds the rows from the statement alone, as PerformanceStudio's desktop viewer does, and the control renders them.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (Viewer4570Tests): the title, the row order, the colour thresholds and the spill flag; they fail on the code before the change.
…rs, rule 38 and the legacy marker (#4585)

Plan analysis: PerformanceStudio's per-rule configuration, rule numbers, rule 38 and the legacy marker. Fixes #4566, refs #4535 and #4530, part of #4511.

- AnalyzerConfig and RulesConfig, ServerMetadata and PlanWarning.RuleNumber come from PerformanceStudio dev, with new Analyze and pipeline overloads that take them; the old overloads use the defaults, so existing findings are unchanged.
- Every rule is guarded by IsRuleDisabled and stamps its RuleNumber at PerformanceStudio's points. ApplySeverityOverrides applies per-rule severity overrides keyed on RuleNumber, skips SQL Server's own warnings and reaches procedure and function bodies (EnumerateAllWithContainer). The configuration source that lets a user set overrides comes next.
- Rule 38: a statement that ran at DOP 2 with a batch-mode operator, without a MAXDOP 2 hint, gets a Standard Edition DOP Limitation finding; it is Info until the server's edition is known.
- The legacy marker: findings of PerformanceStudio's 21 legacy warning types are marked IsLegacy (never SQL Server's own warnings), shown as is_legacy in the MCP plan output and a [legacy] tag in the plan viewer.
- Tests: RuleNumber code and runtime censuses, an all-rules-disabled check, rule 38, the overrides and the legacy marker; a removed RuleNumber stamp fails the census.
Plan viewer: a shared MetricFormatter for costs and durations. Fixes #4577, part of #4511.

- MetricFormatter (PerformanceMonitor.PlanAnalysis) is PerformanceStudio dev's formatter, which its desktop viewer uses.
- The properties panel's estimated costs and the node tooltips' cost rows go through it, so costs show a trimmed, thousands-separated number instead of six raw decimal places; no raw six-decimal cost format is left in the viewer.
- The statement grid's durations go through it in place of a private helper. The sub-millisecond duration format is ported too; the viewer's statement times are whole milliseconds, so it has no caller that shows it yet.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (Viewer4577Tests): PerformanceStudio's 46 formatter cases, and a check that no raw six-decimal cost format remains.
#4591)

Plan viewer: the Wait Stats header's tooltip matches PerformanceStudio. Part of #4511.

- The Wait Stats header in the plan viewer's insights carries PerformanceStudio's desktop tooltip, "{N} ms of waits across {count} wait types", through PlanDisplayText.WaitStatsHeaderTooltip, and clears it when the plan has no waits.
- The header text itself already matched PerformanceStudio's desktop viewer, and the per-card collapse is a web-viewer feature, so neither changes (#4571 is closed as out of scope).
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (Viewer4571Tests): the tooltip text, and no tooltip without waits.
…ection (#4587)

Plan viewer: per-thread stats roll into one collapsed breakdown per section. Fixes #4575, part of #4511.

- The properties panel's per-thread rows for an actual parallel plan collapse into one "Per-thread breakdown (N threads)" expander per section, as in PerformanceStudio dev, with "(skewed: max / min)" in the header when the work is unbalanced.
- ThreadBreakdown (PerformanceMonitor.PlanAnalysis) holds the logic, matching PerformanceStudio: thread 0 is left out, the skew check needs 1,000 rows per worker, a worker share above 0.80 with two workers or 0.50 with more counts as skew, and idle threads are detected.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (ViewerThreadBreakdownTests): the header and skew text, the thread-0 filter, the thresholds and idle threads; they fail on the code before the change, and a threshold mutation fails them.
…bered width (#4588)

Plan viewer: properties panel row model, filter box, copy menu, remembered width. Fixes #4574, part of #4511.

- The properties panel gains PerformanceStudio dev's filter box, a right-click menu with "Copy value", "Copy name and value" and "Copy all properties", and a width that is remembered between plans (default 380, between 280 and 800), as in PerformanceStudio.
- PropertyRows (PerformanceMonitor.PlanAnalysis) holds the filter, the copy text and the width clamp; clipboard writes are guarded.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests: the filter, the copy text for each menu item and the width clamp; they fail on the code before the change.
Plan viewer: add the minimap. Refs #4580 (part 1), part of #4511.

- The plan viewer's toolbar gains a "Minimap" toggle that opens a corner panel with a scaled-down view of the whole plan, ported from PerformanceStudio's desktop viewer: subtree backgrounds in its eight branch colours, a rectangle per operator with the expensive-operator tint and operator icons, elbow connectors, and a box that tracks the scroll position and zoom. Clicking the minimap centres the plan on that point.
- MinimapLayout (PerformanceMonitor.PlanAnalysis) holds the scale, bounds, viewport rectangle, click-to-offset and colour-cycle arithmetic.
- The control is shared, so the Darling Viewer and Lite both get it. Resizing, double-click zoom and accuracy-coloured minimap edges come in part 2.
- Tests (Viewer4580Tests): the layout arithmetic and the colour cycle.
… on a truncated single-statement copy (#4593)

Plan viewer: guard clipboard writes, and hand back the captured query on a truncated single-statement copy. Fixes #4582, part of #4511.

- ClipboardText.TrySetText writes with the same bounded retry as TryRead, and all eight unguarded clipboard writes in PerformanceMonitor.Ui (the plan viewer, the blocking chain, the deadlock graph and the grid export) use it, so a clipboard held by another program no longer crashes the app (PerformanceStudio 7cd9218, 35249e2).
- "Copy Query Text" on a single-statement plan whose statement text was truncated returns the captured query the host already passes to LoadPlan, counted as PerformanceStudio fdd3d22 does.
- The remaining unguarded clipboard writes in Lite and the Darling Viewer outside the shared control are tracked in #4594.
- Tests: the guarded write and the truncated-copy fallback; they fail on the code before the change.
…mn that never clips (#4595)

Plan viewer: PerformanceStudio's wait-row layout, with a benefit column that never clips. Fixes #4576, part of #4511.

- The Wait Stats rows take PerformanceStudio desktop's layout (78a3370): the name and duration columns take the spare width with an ellipsis and a tooltip, the bar and benefit columns size to their content, and each row has a tooltip.
- A new benefit column shows "up to N%" (a whole percent, as in PerformanceStudio's desktop viewer) from the plan's matching "Wait:" finding, through WaitRowText.Benefit.
- The Wait Stats header keeps its tooltip from PlanDisplayText.WaitStatsHeaderTooltip.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (ViewerWaitRowBenefitTests): the benefit text and its join to the wait findings; they fail on the code before the change.
Plan viewer: PerformanceStudio's insights-strip card system. Fixes #4573, part of #4511.

- The insights strip's cards (Runtime Summary, Missing Index Suggestions, Wait Stats) share one neutral surface with a thin per-card accent edge, as in PerformanceStudio desktop 87bad14, instead of a full-colour tinted background, and an empty card goes quiet instead of showing an empty frame.
- InsightCardStyle (PerformanceMonitor.PlanAnalysis) holds the styling rules; the accent brushes are theme resources in the Darling Viewer's and Lite's themes, so no colour is hard-coded in the card code.
- The control is shared, so the Darling Viewer and Lite both get it. PerformanceStudio's Server Context and Parameters cards are tracked in #4597 and #4598.
- Tests (Viewer4573Tests): the card style and the quiet empty state.
…oured edges (#4603)

Plan viewer: the minimap's resize, double-click zoom and accuracy-coloured edges. Fixes #4580, part of #4511.

- The minimap panel can be resized within PerformanceStudio's bounds (160 to 500) and keeps its size across plans; double-clicking a minimap operator zooms to it and selects it; and minimap edges take the actual/estimated row colours from PlanEdgeColour.ForChild, as in PerformanceStudio's desktop viewer.
- MinimapLayout gains the resize clamp, zoom-to-operator and hit-test arithmetic.
- The control is shared, so the Darling Viewer and Lite both get it.
- Tests (Viewer4580Tests): the resize clamp, zoom-to-operator and the hit test; a mutation fails them.
…th blank names on one Azure SQL Database server no longer silence each other (#4908)

Darling and the Viewer matched a whole-server silence by display name, and a blank display name falls back to the host, so silencing one of two blank-named databases on one Azure SQL Database server silenced both, and their alerts shared one dedup key. Rung V157 adds a nullable `config.config_mute_rules.server_id`; a rule with a server id matches only that server's alerts, and a rule without one matches by name exactly as before, so no stored rule changes meaning. The Viewer's Silence, Snooze and "Mute this" write the id; Unsilence also replaces a legacy name-keyed silence for every other server it covered, and only an enabled, unexpired id-keyed silence counts as a replacement. MCP and the web accept and validate `server_id`, store ids may be negative, and the fleet card's `is_silenced` matches by id or legacy name. The alert dedup key appends the store id only when another registration shares the display name, so every other server keeps its key; the MCP `dedup_key` filter matches both forms.
…both apps' grids, MCP reads, slicers, trends and daily summary (#4909)

Blocking and deadlock counts for a time window differed by where you looked: Lite's Overview card and Darling's fleet, Overview and health counts used the event time, while the grids, slicers, trends, severity graphs, MCP reads and daily summary used collection time. Every user-facing windowed read of `blocked_process_reports` and `deadlocks` now windows and buckets on the event's own time (`event_time` / `deadlock_time`). Darling adds `collection_time >= EventWindowFloor.For(start)` with no upper bound, so chunks stay excluded and a late-collected event still counts in the window it happened in; Lite uses the event predicate alone. Hour and day buckets follow the event time. The alert engine's reads stay on collection time (a delivery cursor), as do the DMV snapshot arms and the continuous-aggregate baselines.
… lock-wait line (#4919)

The two-rotation server-log tail test counted exactly one "still waiting" lock-wait line per wait, but PostgreSQL 18 logs that line again when the waiting backend wakes on a latch before the lock is granted, so the test failed intermittently on a loaded runner. The finished marked file and the newest file must now each have at least one line for their wait, with no line read twice (the lines must be distinct by timestamp and text, which a re-read would repeat); the skipped middle file must still contribute nothing, and the files-skipped-by-rotation measurement must still be 1. Tests only.
…yet until it is due (#4915)

On a new server, Lite and the Darling viewer said a collector had never run, or that its data was not collected. They said it before that collector's first run was due. get_running_jobs in both apps said the running_jobs collector was switched off. Now each of these places waits until the collector is due. Until then it says the collector has not run yet. It names when the server started and last collected, and how often the collector runs by default.

- One shared rule decides the first-run grace for every surface in both apps and both MCP tools. A collector is due at the server's first collection, plus its default interval, plus 10 minutes. The rule compares that with the server's last collection. A collector that is off by default gets no grace.
- One shared sentence and one interval formatter give the note. Its times are UTC.
- After the grace passes, the never-ran note comes back.
- The gone-dark and never-ran notes say "usually means" the collector is switched off. running_jobs names its possible cause in one shared text.
- Each surface reads the server's first and last collection with its existing last-run read, so they always come from one source.
…he click, so WPF can't reject it mid-walk (#4920)

Checking or unchecking an item in the Wait Types, Memory Clerks or Perfmon Counters picker could close the Darling Viewer or Lite with "Cannot modify the Visual children for this node because a tree walk is in progress": the handler rebuilt the list containing the checkbox inside its Checked/Unchecked event, which a checkbox's binding can raise while WPF walks the visual tree. A small `PickerRefreshCoalescer` in each app now posts one background-priority callback per burst of toggles; the callback runs the same re-order and chart update, reads the checkbox state when it runs, and applies only when the selected set differs from the one it last applied, so a binding-raised Checked from a rebuilt container can't start another rebuild. Every direct refresh (population, Select All, Clear All, the Perfmon picker's second path) resets that record first. The sort, filter, count text and chart are unchanged.
…ithout collection (#4917)

After Lite archived its data at the 512 MB limit, database state alerts sometimes repeated or were lost, and learned baselines were deleted. The database state editor showed no databases and no overrides. The same happened after 7 days without database state rows for a server. Now the editor, "Reset to current" and the alert sweep read one chosen source, and the sweep gives no verdict when the data cannot be judged.

- The source is the hot table when it holds two or more snapshots for the server. Otherwise it is v_database_states (hot table plus Parquet archive) when that holds two or more. Otherwise it is the hot table.
- GetDatabaseStatesAsync returns null for "no verdict": fewer than 2 snapshots at read time, or a newest snapshot older than 7 days. The alert engine then fires nothing and resolves nothing. An empty list still resolves.
- A stale source changes nothing. A fresh first snapshot still seeds a new server's baselines.
- The sweep chooses the source, maintains the baselines and reads the deviations under one write lock.
- The Darling store never returns null, and its public method stays non-nullable.
- ArchiveService.HotDataDays now holds the 7 days. The archive default, the collection background service and the staleness rule read it.
…ws after the 512 MB archive step (#4921)

After Lite's 512 MB archive step, five features saw only the data collected since that step. They were the reserved-capacity, long-running-jobs and storage-latency advice, High Impact Queries, and the SQL Server Agent check. A stopped Agent then read as one that never ran. The Agent check also lost its running rows 7 days after Agent stopped, when ordinary archival moved them. These reads now use the archive views, which hold the hot rows plus the Parquet archive.

- Reserved capacity reads v_cpu_utilization_stats, long-running jobs read v_running_jobs, and storage latency reads v_file_io_stats.
- High Impact Queries reads v_query_stats for its rows and for both query-text subqueries.
- The Agent status read uses v_agent_status, so ever_seen_running covers every stored row.
- The never-ran probe for database_states reads v_database_states.
- A new test, ArchivableTableBareReadSweepTests, scans Lite's source for bare FROM or JOIN reads of each archivable table. The test fails when a file holds more than its listed count, and each listed read has a reason.
- A second test runs the same scanner over fixed source and checks the exact matches, so the sweep cannot go blind and still pass.
…stem_health event once after a reset stored it twice (#4918)

Before #4887, the first collection after Lite's 512 MB archive step read the collectors' 10-minute fallback window again. It stored blocked process reports, long query completions and system_health events that the archive already held. #4887 stopped new copies, but the stored copies stayed.

Every read of these tables counted such an event twice when both copies fell in its window. Those reads serve blocking counts, alerts, grids and charts, and the daily summary. They also serve analysis facts and baselines, the long query grid and MCP tool, and the system_health parsers. A watermark read that fails also falls back to that window and stores a copy.

- StoredEventCopies is now the one place that reads v_blocked_process_reports, v_long_query_completions and v_system_health_events. Each read passes its own filter, and the helper drops the copies a later batch stored.
- For each identity, every row of the first batch that stored it stays. The helper groups on the identity to find MIN(collection_time) and joins back with IS NOT DISTINCT FROM on every part, so NULL parts match.
- A blocked process report or system_health event is identified by server_id, event_time and a 64-bit hash() of its XML. Two different events that share a server and event time merge only on a hash collision, about 2^-64 per pair. A collision hides one row from one read and never deletes anything. Nothing stores the hash.
- Rows with NULL or empty XML never collapse. A long query completion's identity compares every column exactly.
- Reads look back 600 seconds before their window start (CollectorContext.EventFallbackWindow), so a copy at the window start is still matched to its first batch.
- At 100x volume and a 1 GB memory limit, the window form ran out of memory. The grouped form ran each read in about 4.7 s, close to reading the plain rows.
- A failed watermark read now logs a warning.
…per read, not by a window over the whole archive (#4924)

Lite ran out of DuckDB memory on every alert cycle on a store with a large deadlock archive: the `v_deadlocks` view hid copies of a deadlock with a window over every row's full graph XML, so each read windowed the server's whole history (a 954 MB peak for one read on 30,000 deadlocks). `v_deadlocks` is now a plain union, and copies are hidden per read, after the read's own filter, by one shared identity (server, deadlock time and a hash of the graph; a missing or empty graph or time never collapses). Count and bucket reads count distinct identities; row reads keep one row per identity, the earliest collected with the lowest id breaking a tie; MAX and EXISTS reads use the plain union; the alert read keeps its collected-from bound and look-back. Measured on real data: 49 MB for the biggest read and 275 MB for a nine-server alert cycle, with no out-of-memory error under a 45-read stress and the same counts and rows as before. The write-side duplicate check and the startup cleanup are unchanged.
…he waits that made it fire (#4923)

On Azure SQL Database, the wait-profile anomaly finding named REMOTE_BLOCK_IO first among the waits that led it, although the young-store bar leaves REMOTE_BLOCK_IO out of the rate it tests, so a user chased a wait that didn't make it fire. Both detectors now stamp `bar_excluded_<WAIT>` = 1 beside `contrib_<WAIT>` only when that excluding bar is the arm that fired. `FactAdvice.ComposeAnomalyWaitProfile` lists the waits that counted first and names the excluded one after them as "(not counted toward the threshold on Azure SQL Database)", and `AnomalyIncidentReconciler.DominantWaitFamily` takes the dominant wait from the counted ones unless every contributor is excluded. Both read the marker through one rule, `AnomalyThresholds.IsBarExcluded`. With no marker the text and the reconciler are unchanged, and the reported rate still counts every wait.
…t once and say where the separately monitored databases' events are counted (#4927)

In Lite, an Azure SQL Database logical server's master target counted the blocking events and deadlocks of databases that are also monitored as their own targets: on its Overview card, in the Performance Calendar and the daily summaries, and in its health band. Master's card and daily summary now scope the blocked-process, DMV and deadlock counts with the same separately monitored list and helpers analysis uses, and a deadlock counts on master unless every process in its graph is in the list. Master's Blocking and Deadlocks lists and its Locking & Contention panel keep their rows and say, in one line, that those databases' events are counted under their own servers; Lite's MCP blocking, deadlock and object-locking replies carry the same note. With no list, or when the lookup fails, every read is unchanged and no note shows.
…tart (#4928)

The Darling viewer filled its server list at start, and again only after an add, edit or remove made in that same window. Servers are also added from another viewer, the web viewer and the MCP add and remove tools. A server added there stayed out of the sidebar, the status bar's "Servers: N", the Overview total and the database state editor until a restart. A server removed elsewhere stayed in the list the same way.

- The existing fleet refresh tick reads the config server list and compares its server ids with the ids the viewer loaded. A matching set does nothing more: no reload and no sidebar rebuild.
- When the sets differ, each server that left gets the cleanup a remove in this window gives it. Its favorite pin and alert acknowledgement state are dropped. Both removes now call one shared cleanup method. The list then reloads as it does after a local add or remove, and the selection stays.
- The read is single-flight like the tick's other reads. A failed read is logged and tried again on the next tick. The tick's declared read width goes from 6 to 8.
- Only the config list can change the server list. The load falls back to the observed list, which lacks every configured server that has never collected. The seeded check also reports "not seeded" when it fails, for example on a dead pooled connection after a PostgreSQL restart. So the new GetConfigManagedServersAsync returns nothing unless the store is seeded, and the tick then does nothing. GetManagedServersAsync keeps its fallback for the load. A store that an older service has not seeded keeps today's behavior: changes show after a restart.
- Settings hands the database state editor the viewer's loaded list, so an editor opened after a reload lists the new set.
…nts only its own blocking and deadlocks, and its lists say where its databases' events are counted (#4925)

On a fleet with an Azure SQL Database logical server registered at master and its databases registered on their own, Darling counted the databases' blocking events and deadlocks twice: on master's fleet card, in the header totals, bands and "Needs attention", in MCP get_server_summary, in master's daily summary, the Performance Calendar and the fleet sweep's band, and in the Darling Viewer's fleet totals and per-server summary. Each of those reads now scopes an Azure SQL Database master with the separately monitored list and the scoped SQL alerts and analysis use: a deadlock counts on master unless every process in its graph is in the list. The daily summary filters only its deadlock and blocked-process steps, on every retention tier, and its range cache key carries the list. A server is an Azure master only when its registration and its newest stored engine edition both say so, on every surface. A failed lookup or scoped read keeps that master's unscoped counts with a warning, and the rest of the fleet read succeeds. Master's Blocking and Deadlocks lists and its Locking & Contention panel keep their rows and say, in one line, that those databases' events are counted under their own servers; MCP get_blocking, get_deadlocks and get_object_locking return the same note. Every other server, and a master with no separately monitored databases, runs the same SQL as before.
…y name (#4926)

Alert History listed analysis alerts under the server's internal storage name (host:database) while engine alerts used the display name, so filtering Alert History by a server's display name hid every analysis alert for that server and the Server column showed two spellings for one server. The web page, MCP get_alert_history, the triage page, the Darling Viewer and Lite now show each row under the server's current display name: Darling's alert reads join servers and return COALESCE(display_name, server_name), and Lite maps each row's server id to its display name at read time. The alert notebook's resolution and re-fire reads use the same aliased join. Each row keeps its stored spelling for "Mute this" and the tray Snooze, so mute rules still match. What is stored, the analysis bucket and repeat-alert keys, and dismiss are unchanged. When two servers share a display name, an analysis row's triage link opens without a server scope, as engine rows already did.
…d inserting it again (#4929)

The persistence gate's record is one INSERT ... ON CONFLICT DO UPDATE. The incident-occurrence set is upserted per fingerprint, then the metric's other fingerprints are deleted, in one transaction, so a surviving key is never deleted and inserted again. Comments in Schema.cs and QueryStoreSliceRepairService.cs now describe what DuckDB 1.5.5 does.
…ser's list once (#4931)

#4884 added RBIO_COMM_RETRY and SQP_STATS_REPORTING to Lite's default ignored waits, but Lite reads the per-user ignored_wait_types.json first and the bundled file is copied only when that one is missing, so an install upgraded from an earlier release never received the two names and kept collecting, showing and analysing them. At startup, before anything loads the list, Lite now merges the bundled defaults the per-user file has not seen into it once: the file gains a seen_defaults list, and a file without one is taken to have seen v3.8.0's list (identical in every release through v3.8.0), so an upgraded install gains exactly the defaults added since, even after skipping releases. The user's removals stay removed, their additions, order and other settings stay, and the file is rewritten only when something changed, atomically, keeping its line endings. A malformed or unexpected file is left exactly as it was, with a warning, and startup never fails. Darling uses compiled-in defaults and needs no change.
…ert sweep after a start (#4933)

* Lite: an Azure SQL Database master target is scoped from the first alert sweep after a start

The master scope read the engine edition from the live connection status only. The first sweep after a start runs before any connection check has read it, and a failed check or an edit blanks the status, so in those sweeps a master target counted the blocking and deadlocks of the databases that alert on their own targets.

KnownEngineEditions now gives the alert sweep and the analysis provider one rule: a live edition wins and is remembered; without one, the stored edition, seeded in one grouped read before anything that reads the scope starts; with neither, unscoped as before.

* Lite master scope: the startup seed of the stored engine editions waits at most 10 seconds

Collection, the alert engine, MCP and the server list start after the seed, so a read that stalls
(a held lock, a slow Parquet scan) no longer holds them up. A read past the limit is logged as a
warning and the start goes on. A late result still seeds, and never replaces an edition that a live
status reported meanwhile. A late failure is logged, so its exception is always observed.
…wn databases, and the Viewer card shows no sibling's Last (#4932)

#4925 scoped an Azure SQL Database master's blocking and deadlock counts to its own databases but left the last-seen times alone, so a master whose own databases had no deadlocks showed "0 deadlocks, last seen minutes ago" on the fleet card from a separately monitored database's deadlock, and the Darling Viewer's Overview card read "Last: N ago" from a sibling's event under 0/0. The fleet card's deadlock_last_seen now comes from the same single pass as the scoped count, by the count's own rule, so the two can't disagree and a failure keeps that master's whole unscoped row; get_server_summary makes one graph pass instead of two. In the Viewer, a master with separately monitored databases shows no "Last" on its Overview card for blocking or deadlocks, because an own-database newest over all stored history would need an all-history read or a cache on every refresh; its counts stay scoped and its lists keep every row with their note. Every other server keeps its "Last".
…indexes at every open (#4930)

* Lite reopens its database after a fatal DuckDB error and repairs its indexes at every open

A fatal DuckDB error invalidates the whole database: every later statement fails until every
connection closes and the file opens again, so collection and alerting stopped for every server
until Lite was restarted. Lite now reopens the same file under the write lock (3 attempts with
back-off), says so in the status bar and on Collection Health, and logs the fatal error and the
reopen.

At every open, an explicit CHECKPOINT and then a rebuild of each explicit index (each DROP and
CREATE committed on its own) keep rows that WAL replay restored in their indexes (duckdb#26106),
so a later delete over them no longer fails with a FATAL error.

* Lite opens an existing file when a declared index cannot be created

The index repair at the open can drop an index that then cannot be built again, for example one
too large to build within the memory limit. The schema's CREATE INDEX IF NOT EXISTS statements then
met the same failure and stopped the start, and every start and reopen attempt after it. On an
existing file, a declared index that cannot be created now logs one ERROR and the start carries on,
so the next start tries again (the same pattern as the missing-column heal). A fresh file still
fails, because there a declared index that cannot be built is a bug.

Also: the comment no longer names a server removal as a delete over indexed rows (it deletes none).

* LocalTime takes a block body, so the source member scan reads it whole

TsqlConventionGuardTests.TheMemberScan_ReadsEveryDeclarationWhole reads a multi-line expression-bodied member short of its end. The same expression in a block body reads whole, so no entry is added to KnownTruncatedRanges.

* Lite reopens its database only once running collections end, and pauses collection while it is down

A reopen attempt takes the collection gate the size-triggered reset uses, so it runs after every registered collection has ended and none starts until it is done. Collector runs, the Query Store backfill and the loop's housekeeping skip while the database is down, a per-database loop ends at the first fatal error, and the backfill's connections take the database lock. A fatal error in the open's CHECKPOINT, index rebuild or declared index statements stops the open with an error that names the step. One run reopens at most five times. A dispose during an attempt closes the sentinel the attempt opened. An existing file is told by whether it was there before the open.

* A collector run that starts while Lite's database is down skips, and one line says so each way

* A reopen that waits long for running collections warns, and the paused and resumes lines are written in order

Past the collection gate's 3-minute drain timeout, the wait for running collections logs a Warning with their count, and again after each further 3 minutes. It keeps waiting. The health read, the compare and the paused or resumes line in LocalDatabaseIsDown are one step under a lock, so a caller holding a read from before a reopen cannot log the change backwards.
…efers, and a start after an interrupted update puts the store's own runtime back (#4934)

When a package ships a new PostgreSQL runtime, the update clears the last update's rescued runtime and renames the live runtime into it; just after the store stops, an antivirus scan or the exiting server can hold that folder for a moment, and either step gave up at once, leaving the host on the old runtime until the service next started. Both steps now retry IO and access errors up to five times over about 5.5 seconds, and only then fall back exactly as before. The revert after a failed extract logs why the extract failed and retries its two moves the same way, so a short lock no longer leaves no runtime at all. A start that finds the runtime missing after an interrupted update now puts back the runtime that last opened the store and retries the update, instead of extracting the new one as a first run; it acts only for an interrupted same-major update (a readable stamp that differs from the shipped package, a rescued runtime of the store's major that carries its TimescaleDB versions, no server running on the store), and in every other case the start re-extracts the package as before.
…e store, and an unfinished update says how to recover (#4935)

A runtime update moves the live runtime aside into pg-runtime-prev before it extracts the new one, and until the new runtime is known to open the store that copy is the only one that does; nothing recorded that, so a failed extract that left a partial runtime holding pg_ctl.exe let the next start clear pg-runtime-prev and delete the only good runtime. A rescue marker in pg-runtime-prev now exists exactly while it holds the only runtime known to open the store: it is written before the live runtime is moved aside and removed wherever the runtime is known to open the store again. While it exists, pg-runtime-prev is cleared only when it provably cannot open the store or the update that wrote the marker provably finished; anything else, a failed version probe included, defers the update, logged with the marker file to delete, and a start that then fails says the same at Error. Under the marker the start-time restore puts the rescued runtime back even over a partial runtime, and a stamp cut short to zero length no longer stalls updates. A store with no TimescaleDB record is taken to be on 2.28.1 outside the marker.
The 3.9.0 release, cut from dev 9e1b619. The version goes from 3.8.0 to 3.9.0. In the CHANGELOG, the 09-28 to 09-30 merges and the managed store's random_page_cost entry are spliced into [Unreleased], which then becomes [3.9.0] - 2026-10-02 with an Important section (collectors on large fleets run at their configured interval; upgrading from 3.8 runs store migrations V126 through V157 on the service's first start), the 45 Fixed and 1 Changed entries of the fixes merged since the last splice with their link definitions, and two amended entries (#3752, #3691). The 3.9.0 entries move to docs/changelog/3.9.md with a regenerated archive census, and the migration upgrade test gains the v3.8.0 schema ladder it climbs from. No product code or configuration changes.
…t name (#4939)

Wait rows stored before the collectors trimmed wait names keep the DMV's
trailing space, for example "SQP_STATS_REPORTING ", until retention removes
them. Lite compared the stored name exactly, so the ignore list never hid the
spaced spelling, and a wait stored both ways came back as two rows.

Every Lite read that filters, groups or looks up by wait name now compares
rtrim(wait_type): the ignore list hides both spellings, the two spellings sum
into one row under the clean name, and a lookup by the clean name finds the
spaced history. This covers the Wait Stats grid, picker and trends, the
overview total wait line, Current Waits, the wait drill-down, the MCP wait
tools, the daily summary's top wait, the FinOps wait categories, and the
analysis wait facts and anomaly contributors.
… up to 1,000 files (#4940)

The Darling PG and Linux jobs' gate-notice steps received the whole changed-file list in one environment variable, and Linux refuses to start a process whose single environment string passes 128 KiB, so the 3.9.0 dev-to-main pull request (2,843 files) failed the Linux job at that step with "Argument list too long". Both steps now pass each list only when its count is 1,000 or fewer, with a short placeholder otherwise; the counts always print and the path-filter step still logs every file. DarlingPathFilterGateTests now pins the cap expression on both jobs.
…s clean name (#4941)

* Darling reads a wait stored under both spellings as one wait under its clean name

SQL Server reports four wait names with a trailing space, and from 3.9 the collector
stores them trimmed (#4884). History from before the upgrade keeps the spaced name,
so the same wait could be stored under two spellings.

Reads that group by wait now key on rtrim(wait_type): the viewer's wait picker, wait
trends, current-waits trend and FinOps category summary, the MCP wait stats, wait
types and current-waits trend, the daily summary's top wait, the analysis wait
facts and the anomaly wait contributors. Lookups by name keep the column bare and
match both spellings: the MCP wait trend and the viewer's queries-by-wait drill-down
use wait_type IN ($n, $n || ' '). Custom Views group the three SQL Server wait-name
dimensions on the trimmed name and widen each filter value instead of wrapping the
column. pg_wait_stats (PostgreSQL wait events) is unchanged.

* Wide wait_stats grouping reads sum per stored name, then merge the spellings on rtrim

rtrim in GROUP BY ran once per row and cost 22-30% on the 7-day wait stats
and 10-day daily summary reads. The six wide grouping reads now keep the old
per-name aggregation inside and merge the two spellings over those groups,
which measures within a few percent of the read before #4941.

The FinOps wait categories read the category from the clean name.

Tests: gte/lt guard cases and a top-N time series with "(other)" in the
live spaced-history class; the SQL text tests follow the two-level shape.
The 3.9.0 entry for the wait-name fix (#4884) now covers #4939 and #4941, which completed it after the release was cut: its references gain both PRs in the index and the archive, rows stored with the trailing space before the upgrade are described as reading as the clean name in both apps, and the entry says Lite's wait views hide the two newly ignored waits in history from before the upgrade while Darling shows that history until retention removes it. The archive census carries the new 3.9.0 prose hash; the entry count is unchanged.
The five SignPath signing steps in the release build (Sign Lite, Sign Darling, Sign Lite (Velopack), Sign Darling Viewer (Velopack) and Sign the executables Velopack generates) move from signpath/github-action-submit-signing-request@v2 to @V3, which sends signing requests to SignPath's Pipeline Connector. The connector needs the SignPath GitHub App, installed on this repository since 2026-09-24, and a comment above the first step says so and notes that SignPath signs only a run's first three attempts. No other step input changes; signing runs only when a release is published. SignPathActionVersionTests pins exactly five signing steps, all on v3.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 2, 2026 11:44
@erikdarlingdata
erikdarlingdata merged commit 4554345 into main Oct 2, 2026
28 of 29 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant