Skip to content

Make the read-lock cancellation flake deterministic under a loaded runner - #4655

Merged
erikdarlingdata merged 3 commits into
devfrom
feature/analysispass-readlock-flake
Sep 28, 2026
Merged

erikdarlingdata merged 3 commits into
devfrom
feature/analysispass-readlock-flake

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Test-only flake fix. No issue: this addresses the flake in Lite.Tests/AnalysisPassTokenThreadingTests.cs:301, TheReadLockWaitIsAbandonableWhileAWriterHoldsIt, seen failing alone in a slow CI build on 2026-09-12 and again in a loaded local full Lite run on 2026-09-28.

Why

The test guards a real guarantee: a store read waiting on DuckDbInitializer's read lock can be abandoned through its CancellationToken while a writer holds the lock, instead of blocking for the writer's full hold. That guarantee is real and unchanged by this PR.

What flaked was how the test triggered the cancellation. It armed new CancellationTokenSource(TimeSpan.FromMilliseconds(250)) and trusted that 250ms was long enough for the read-lock attempt to already be polling by the time the token fired. A delay-based CancellationTokenSource cancels through a System.Threading.Timer callback, and that callback has to be scheduled by the CLR ThreadPool. Under a loaded runner two things push that scheduling past the 250ms mark, and past the test's 10-second backstop:

  • A parallel run of the rest of the suite queues real work onto the same process-wide ThreadPool the timer callback needs a slot on.
  • DuckDbInitializer's s_dbLock is one static field, shared by every DuckDbInitializer in the process. A parallel test class that also touches it (there are well over a hundred in Lite.Tests) creates real, legitimate queuing on the same lock, which also affects how long the writer itself takes to acquire it (bumped from a 5s to a 30s wait below, for the same reason).

The lock's own poll was never the slow part. AcquireReadLock(CancellationToken) (Lite/Database/DuckDbInitializer.cs) only checks its token between 50ms TryEnterReadLock attempts, so it reacts within about one interval of the token actually firing. The problem was getting the token to fire promptly at all under contention - confirmed by reading the implementation, not by directly capturing a failure message. I could not reproduce the failure live (see Test plan): 30 solo runs under 28 CPU-bound background processes on a 32-core box passed clean, and a further ~8-9 runs of a 9-class combination chosen for AcquireReadLock/AcquireWriteLock/LocalDataService contention, on a machine already busy with other builds and test runs, also passed clean. A CancellationTokenSource-Timer-vs-ThreadPool race is a genuine race, not a guaranteed reproduction on demand, so the fix is grounded in reading System.Threading.CancellationTokenSource's and DuckDbInitializer's actual code paths.

I did not find a product bug. The lock implementation's cancellation handling is correct as written; this is entirely about how the test manufactured its cancellation.

What changes

TheReadLockWaitIsAbandonableWhileAWriterHoldsIt in Lite.Tests/AnalysisPassTokenThreadingTests.cs:

  • Replaced the timed CancellationTokenSource(250ms) with an untimed one, cancelled explicitly by the test thread via cts.Cancel() - no Timer, no ThreadPool callback in the critical path.
  • The read-lock attempt now runs on its own Task.Run (no token passed to Task.Run itself, only to AcquireReadLock, so a thrown OperationCanceledException faults the task rather than putting it in the Canceled state - awaiting it rethrows the same exception type instead of a wrapped TaskCanceledException).
  • A new readerIsWaiting signal is set immediately before that task calls AcquireReadLock, and the test waits for it before proceeding toward Cancel(). A second commit on this same PR (below) found that this signal alone was not sufficient and tightened it further.
  • The writer's own readiness wait went from 5s to 30s, matching the other bounds already in the test, because s_dbLock being static means a parallel suite run can legitimately queue the writer behind another test's hold of the same lock.
  • The one wall-clock number left is a 10-second bound on reader.WaitAsync(...), used only as a hang backstop (documented as such in the test's own comment) - a working cancellation returns in about one 50ms poll interval, so this only matters if the lock genuinely stopped observing its token.
  • Dropped the now-unused Stopwatch/System.Diagnostics import.

Checking for the same timing shape elsewhere: no other test in this file has a CancellationTokenSource/Task.Run/Wait(TimeSpan) shape - the rest of the class is static source-regex scanning or pure in-memory logic. The Darling twin (Darling/Darling.Tests/AnalysisPassTokenThreadingTests.cs) has no read-lock timing test at all; its own doc comment already states this guard is "the Lite-only half of #2443" because Darling has no equivalent static file lock. git grep -l AnalysisPassTokenThreading turns up two more files (Lite.Tests/CrossAppGuardCiGateTests.cs, Darling/Darling.Tests/CommentFilterAdoptionTests.cs) but both only reference this file's path in a bounded reachability/doc-comment-scanning list, not a copy of the timing shape - confirmed by reading both entries, nothing to fix there.

No product code changed - only the test file.

Second pass: the reader-waiting signal alone was a false green

A second look at this test found a gap: readerIsWaiting.Set() runs as the very first statement inside the reader's Task.Run, before it calls AcquireReadLock at all. The test then does readerIsWaiting.Wait(...) and calls cts.Cancel() immediately afterward. That proves the reader task has started - it does not prove the reader has reached, let alone blocked in, AcquireReadLock's poll loop. In practice cts.Cancel() almost always wins that race and lands before the reader's first ThrowIfCancellationRequested() check even runs (see Proof below - it was 10/10 in testing, not an occasional race).

That matters because the test's own name promises the wait is abandonable while a writer holds the lock - cancellation reaching a reader that is genuinely blocked, not a reader that sees an already-cancelled token before it ever tries to enter the lock. The real AcquireReadLock (Lite/Database/DuckDbInitializer.cs:194-208) checks the token before each poll, so it happens to pass either way. But the most natural regression - "check the token once, then block uncancellably" (cancellationToken.ThrowIfCancellationRequested(); s_dbLock.EnterReadLock(); in place of the polling loop) - would ALSO pass the test as it stood after the first round, purely because the cancel lands before that single check. The old 250ms-timer version of this test caught that same regression shape, essentially by accident, via the timer's own scheduling delay giving the reader time to actually block; tightening the timer away without closing this gap would have quietly dropped that coverage.

Fix: after readerIsWaiting.Wait(...), the test now also spins (SpinWait.SpinUntil, 30s backstop) on DuckDbInitializer.s_dbLock's own WaitingReadCount (read via reflection, since the field is private static) until it shows a thread genuinely blocked trying to enter the lock, then waits a 100ms settle margin (two poll intervals; nothing is asserted on it), and only then calls cts.Cancel(). s_dbLock is one static field shared by every DuckDbInitializer in the process, so a parallel run of the rest of the suite could in principle tick WaitingReadCount over 1 via another test's own blocked reader - this paragraph originally claimed that was harmless, since this test's own writer still holds the lock regardless of what tripped the spin. That claim was wrong, and "Third pass" below describes the actual gap and its fix. The class's XML doc comment previously asserted a guarantee ("the reader is guaranteed to be genuinely waiting... when the cancellation arrives") that the two signals alone did not actually provide; it was corrected in this commit to credit the WaitingReadCount spin instead, and is corrected again, more precisely, by the third pass below.

Proof

  1. False green, reproduced. Planted a scratch-only regression in Lite/Database/DuckDbInitializer.cs's AcquireReadLock else branch: cancellationToken.ThrowIfCancellationRequested(); s_dbLock.EnterReadLock(); in place of the while (!TryEnterReadLock(...)) polling loop - i.e., check once, then block uncancellably. Checked out the PR's then-current test (pre-tightening, git checkout origin/feature/analysispass-readlock-flake -- Lite.Tests/AnalysisPassTokenThreadingTests.cs) against that plant and ran it 10 times: 10/10 passed, each in ~0.14-0.18s (Total: 1, Errors: 0, Failed: 0) - the false green, and a near-certain one in practice rather than a rare race: cts.Cancel() on the test thread reliably outran the reader task reaching its first (and, under the plant, only) token check.
  2. Tightened test catches it. Restored the tightened test against the same plant and ran it twice: 2/2 failed, both at ~10.3-10.4s, with Assert.ThrowsAny() Failure: Expected: typeof(System.OperationCanceledException); Actual: typeof(System.TimeoutException). The failure landed at the Assert.ThrowsAnyAsync line, not at the earlier SpinWait.SpinUntil assertion ("no reader ever blocked on the lock"), which confirms the spin itself resolved correctly - the reader really was blocked in the uncancellable EnterReadLock() when Cancel() ran, it just couldn't observe the cancellation, and the 10s hang backstop caught that.
  3. Original plant re-checked. Swapped the plant to the token-ignored-entirely shape from the first round's own proof (bare s_dbLock.EnterReadLock(), no token check at all, regardless of CanBeCanceled). Tightened test, run once: still red, same TimeoutException shape - confirming the tightening didn't narrow coverage of the regression the first round was already built to catch.
  4. Plants discarded. git restore Lite/Database/DuckDbInitializer.cs; git diff origin/dev -- Lite/Database/DuckDbInitializer.cs is empty and git status --short is clean - no product code in this PR.
  5. GREEN under load. Rebuilt against the real, unmodified product code. Ran the class 30 times in a row while 28 CPU-bound background processes spun on a 32-core box (matching the first round's own load methodology), then killed those 28 processes and confirmed none were left running: 30/30 passed, Total: 6, Errors: 0, Failed: 0 every run, 0.79-1.05s each (vs. ~0.4-0.5s unloaded).
  6. Full suite. Merged origin/dev (already up to date, no conflicts). Full Lite.Tests suite, once: Total: 5525, Errors: 0, Failed: 0, Skipped: 0, Not Run: 0, Time: 209.425s.

Third pass: WaitingReadCount alone still let a different test class stand in for this one

The same fix had a second gap. Lite.Tests runs its test classes in parallel - there is no CollectionBehavior, no runner JSON serializing classes, and no [Collection] attribute on AnalysisPassTokenThreadingTests - and roughly twenty other analysis test files in the suite call AcquireReadLock on the same process-wide static s_dbLock. The second pass's spin only checked dbLock.WaitingReadCount >= 1, with no way to tell whether the thread that tripped it was this test's own reader or some other class's reader blocking on the same lock at the same moment. If another class's reader ticks the count first, while this test's own reader task is still sitting on a ThreadPool queue and has not yet reached its own AcquireReadLock call, cts.Cancel() fires before this reader's own token check ever runs. A "check once, then block uncancellably" regression passes anyway in that interleaving, since ThrowIfCancellationRequested() sees an already-cancelled token on its one and only look - a false green, in the same shape the second pass had already found and thought it had closed. The 100ms settle after the spin narrowed the window but could not close it: it is a fixed margin against a scheduling delay of unknown size, not a check tied to this reader's own state.

Fix: the test now captures this reader's own Thread object from inside its Task.Run delegate (published safely across threads by the readerIsWaiting.Set()/.Wait() signal that already separates the two threads) and spins on dbLock.WaitingReadCount >= 1 && (readerThread.ThreadState & ThreadState.WaitSleepJoin) != 0 instead of the count alone. WaitSleepJoin on THIS reader's own thread can only be true after it has made its own check - real or regressed - and is now genuinely blocked inside TryEnterReadLock or EnterReadLock; nothing else on that thread waits between Set() and the lock. That closes the interleaving above: by the time Cancel() runs, this reader is either past its check and correctly still polling (real implementation, still observes the cancellation within one ReadLockPollInterval) or past its one check and blocked uncancellably (regressed implementation, now provably fails instead of coincidentally passing). The now-redundant 100ms settle delay is dropped. The class doc comment, the inline comment above the spin, and the "cannot turn a real failure into a false pass" claim in "Second pass" above are corrected to describe this actual guarantee instead of the one that turned out not to hold.

Proof

  1. Regression shape still caught. Planted the same "check once, then block" regression as the second pass (cancellationToken.ThrowIfCancellationRequested(); s_dbLock.EnterReadLock(); in AcquireReadLock's else branch), rebuilt, ran the third-pass test 3 times: 3/3 failed, each at ~10.5-10.6s, Assert.ThrowsAny() Failure: Expected: typeof(System.OperationCanceledException); Actual: typeof(System.TimeoutException) - the fix still catches this regression.
  2. Token-ignored-entirely, re-checked. Swapped to the token-ignored-entirely plant (bare s_dbLock.EnterReadLock(), no check at all, regardless of CanBeCanceled) from the first round's proof, ran 3 times: 3/3 failed, same TimeoutException shape.
  3. Plants discarded. git checkout -- Lite/Database/DuckDbInitializer.cs; git status --porcelain clean except this test file. git diff origin/dev -- Lite/ shows only Lite/Themes/CoolBreezeTheme.xaml differing, which is dev's own forward drift and not this branch's: git diff <merge-base> HEAD -- Lite/ (this branch's own commits since the merge base with origin/dev) is empty, confirming no product file is in this PR.
  4. GREEN under load. Started 4 CPU-bound background processes (yes > /dev/null, independent of this shell, PIDs confirmed live via ps -W), ran the class 30 times in a row: 30/30 passed, Total: 6, Errors: 0, Failed: 0 every run, 0.47-0.67s each. Killed all 4 by PID afterward and confirmed none were left running before continuing.
  5. Full suite, once. Rebuilt clean, ran the full Lite.Tests suite once (not under the synthetic load - the suite's own ~5500 tests across many parallel classes already provide it): Total: 5525, Errors: 0, Failed: 1, Skipped: 0, Not Run: 0, Time: 313.290s. The one failure, StatusBarSizeReadLockTests.GetUsedDataSizeMb_WhenTheWriteLockIsHeld_GivesUpInsteadOfBlocking, is a separate timing flake in a file this PR does not touch, and is not evidence of a regression from this change.

Test plan

  • dotnet build Lite.Tests/Lite.Tests.csproj -c Debug - Build succeeded, 0 Warning(s), 0 Error(s) (confirmed across all three rounds and every plant/revert cycle).
  • RED proof (round 1): planted a scratch-only regression in Lite/Database/DuckDbInitializer.cs (AcquireReadLock calling bare s_dbLock.EnterReadLock(), ignoring the token entirely - the pre-The analysis token is armed but 167 Darling and ~138 Lite store calls never receive it #2443 shape), rebuilt, ran the single test. It failed deterministically on the first run: Assert.ThrowsAny() Failure: Expected: OperationCanceledException, Actual: TimeoutException, at ~10.3s (the hang backstop doing its job). Reverted the plant immediately after; git diff --stat against Lite/Database/DuckDbInitializer.cs is empty, confirming no product code is in this PR.
  • FALSE GREEN proof (round 2): see "Proof" above, items 1-4.
  • GREEN, class alone, unloaded: 6/6 pass in under 1s.
  • GREEN, class only, 30 runs in a loop under 28 CPU-bound background processes on a 32-core machine: 30/30 clean, each run 0.79-1.05s (see "Proof" above, item 5).
  • Additional load probe (not part of the 30 runs above): ~8-9 runs of a 9-class combination (this test plus every Lite.Tests class that touches AcquireWriteLock/AcquireReadLock/LocalDataService directly) on an already busy machine - clean, stopped early because each run took ~30s.
  • Merged origin/dev before the final run (round 1: clean, no conflicts; round 2: already up to date).
  • Full Lite.Tests suite, once, after the round-2 merge: Total: 5525, Errors: 0, Failed: 0, Skipped: 0, Not Run: 0, Time: 209.425s.
  • RED proof (round 3): planted "check once, then block" regression, ran the third-pass test 3x - 3/3 failed (TimeoutException at ~10.5-10.6s). Planted "token ignored entirely" regression, ran 3x - 3/3 failed, same shape. Reverted both; git diff <merge-base> HEAD -- Lite/ empty, confirming no product file in this PR (see "Third pass" Proof above, items 1-3).
  • GREEN, class only, 30 runs in a loop under 4 CPU-bound background processes: 30/30 clean, 0.47-0.67s each; processes killed by PID afterward (see "Third pass" Proof above, item 4).
  • Full Lite.Tests suite, once, after the round-3 fix: Total: 5525, Errors: 0, Failed: 1, Skipped: 0, Not Run: 0, Time: 313.290s - the one failure is in StatusBarSizeReadLockTests.cs, a file this PR does not touch (see "Third pass" Proof above, item 5).
  • Installer.Tests not run (out of scope). AvailabilityGroupsTabRefreshTests.cs, StatusBarSizeReadLockTests.cs, and the theme tests are not touched by this PR.

CHANGELOG

SECTION: None - test-only change; no user-visible behavior changes.

…nner

TheReadLockWaitIsAbandonableWhileAWriterHoldsIt armed a
CancellationTokenSource(250ms) and trusted that the delay was long
enough for the read-lock poll to already be running before it fired.
That delay cancels through a Timer callback, which the CLR ThreadPool
has to schedule; under a loaded runner (a parallel run of the rest of
the suite queues real work onto the same pool, and DuckDbInitializer's
lock is one static field every test in the process contends for) that
scheduling can lag the 250ms mark by seconds, past the test's 10s
backstop. The lock's own poll was never the slow part - it only checks
its token between 50ms TryEnterReadLock attempts.

The reader now cancels explicitly instead of on a timer: the writer
signals once it holds the lock, the reader signals once it is about
to wait, and only once both are observed does the test call Cancel()
itself. The writer's own wait is bumped from 5s to 30s for the same
reason - the lock is static and process-wide, so a parallel run can
legitimately queue it behind another test's hold.
… cancelling

readerIsWaiting only proved the reader Task had started, not that it had
reached AcquireReadLock's poll loop yet. Cancelling right on that signal
usually raced ahead of the call and threw out of the loop's very first
ThrowIfCancellationRequested() check - which a "check once, then block
uncancellably" regression in AcquireReadLock would pass too, since the
token is still uncancelled at that first check. Spin on s_dbLock's own
WaitingReadCount instead, which only turns positive once a thread is
genuinely blocked entering the lock, then cancel. Also corrects the
class doc comment, which claimed the two ManualResetEventSlim signals
alone already guaranteed the reader was genuinely waiting when Cancel()
ran; that guarantee didn't exist until this change.
Lite.Tests runs its classes in parallel with no [Collection] to
serialize them, and roughly twenty analysis test files besides this
one call AcquireReadLock on the same static s_dbLock. That let
another class's reader trip WaitingReadCount >= 1 while this test's
own reader had not yet reached AcquireReadLock, so Cancel() could
land before this reader's first token check - a false green for a
"check once, then block" regression in that interleaving, despite
the class doc comment's claim that this could not happen.

Capture this reader's own Thread from inside its Task.Run and spin
on WaitingReadCount >= 1 together with that thread's own
ThreadState.WaitSleepJoin. That state is only reachable after this
reader has made its own check (real or regressed) and is genuinely
blocked inside TryEnterReadLock or EnterReadLock, so Cancel() can no
longer land early. Drop the now-redundant 100ms settle delay, and
correct the class doc comment, the inline comments, and the PR
body's "cannot turn a real failure into a false pass" claim to
describe the actual guarantee.

Proof: planted "check once, then block" and "token ignored
entirely" regressions in AcquireReadLock; the hardened test goes red
(TimeoutException at the 10s backstop) 3/3 for each, reverting
cleanly with no product file touched. 30/30 green for the class
under CPU load, and the full Lite suite passes except one
pre-existing failure in a different lane's file.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 28, 2026 22:24
@erikdarlingdata
erikdarlingdata merged commit 2d44c9b into dev Sep 28, 2026
16 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/analysispass-readlock-flake branch September 28, 2026 22:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant