Watchdog the macOS test-host hang from outside the process - #427
Conversation
PlanViewer.Core.Tests can wedge in a CoreCLR GC-suspension livelock on macOS ARM64
(.NET 10.0.8) and spin a full core until someone notices the heat. It happened
twice unprompted today: three hosts at ~108% CPU each, one of them wedged 22
minutes and still going when sampled. Root-caused from a sample: SuspendEE present
in all 2451 samples, the victim thread interrupted where CheckActivationSafePoint
can never succeed, and the mach-exception thread churning thread_get_state /
thread_set_state forever.
The load-bearing fact is that nothing inside the test host can stop it. xUnit
timeouts and in-test CancelAfter cannot fire because the execution engine itself is
suspended - that IS the livelock - and dotnet-stack and EventPipe hang for the same
reason. Only an out-of-process watchdog works, and both mechanisms here are
enforced by vstest.console, a separate process from the host it watches.
Two layers, deliberately:
- The three CI invocations (ci, nightly, release) now pass the blame-hang options,
which is what produces a sequence file naming the test that wedged.
- RunSettingsFilePath in the test csproj points at hang-watchdog.runsettings, so a
bare `dotnet test` with no flags gets a TestSessionTimeout backstop. This is the
layer that matters most: BOTH incidents were plain local runs, so a fix living
only in the workflow files would protect the machine that was never at risk and
leave the laptop unprotected.
15 minutes is far above any real runtime (individual tests are milliseconds, the
whole suite normally well under a minute), so it only fires on a genuine wedge.
Verified rather than assumed, because a silently-ignored runsettings would look
exactly like a working one: temporarily set TestSessionTimeout to 1ms and a bare
`dotnet test` reported "Aborting test run: test run timeout of 1 milliseconds
exceeded", then restored it. The blame collector confirms itself active on a normal
run ("All tests finished running, Sequence file will not be generated"). Zero test
hosts survived any run.
Two XML-comment errors on the way in, both caught by the toolchain rather than by
me: '--' is illegal inside an XML comment, and I had written the flag names and a
dash-separated aside into comments in both the csproj and the runsettings.
Evidence for a future dotnet/runtime report is preserved outside the session
scratchpad at ~/Documents/dotnet-hang-evidence (sample, gzip, and a README
explaining the signature); related to dotnet/runtime#66759.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Merging on While checking it I found the review gate has been skipping repo-wide since 2026-08-03 and reporting success while doing it — filed as #428. Root cause is mine: #423 merged the Dependabot-skip into This PR would have skipped review regardless, since it modifies three workflow files and that always fails the identity check by design. What I'm relying on instead: Merging because it stops a recurring problem that has burned three cores twice today and is currently detected by noticing the laptop is hot. |
What does this PR do?
Stops
PlanViewer.Core.Testsfrom spinning a full core indefinitely when it wedges on macOS ARM64.It happened twice unprompted today: three test hosts at ~108% CPU each, one wedged 22 minutes and still going when sampled. It was noticed because the laptop got hot, which is not a monitoring strategy.
Why nothing in the test host can fix this
This is a CoreCLR GC-suspension livelock (.NET 10.0.8, macOS ARM64), not a slow or hanging test. From a 1 ms
sampleof a wedged host:GCToEEInterface::SuspendEE→ThreadSuspend::SuspendEEin all 2451 samples — the runtime is trying to suspend for GC and never completes.CheckActivationSafePointcan never succeed, so the runtime re-sends activations forever.thread_get_state/thread_set_state, which is what burns the core.So xUnit timeouts and in-test
CancelAftercannot fire — the execution engine itself is suspended; that is the livelock.dotnet-stackand EventPipe hang for the same reason. Only an out-of-process watchdog works, and both mechanisms below are enforced byvstest.console, which is a separate process from the host it watches.Two layers, and the second is the one that matters
ci,nightly,release) pass the blame-hang options, which produce a sequence file naming the test that wedged.RunSettingsFilePathin the test csproj points athang-watchdog.runsettings, giving a baredotnet testaTestSessionTimeoutbackstop with no flags to remember.Layer 2 is the important one: both incidents were plain local runs. A fix living only in the workflow files would protect the machine that was never at risk and leave the laptop unprotected — which is precisely how this recurred after being root-caused the first time.
15 minutes is far above any real runtime here (individual tests are milliseconds; the suite is normally well under a minute), so it only fires on a genuine wedge. It's a backstop, not a performance budget.
How was this tested?
A silently-ignored runsettings looks exactly like a working one, so I verified the mechanism rather than the file's existence:
TestSessionTimeoutto1, ran a baredotnet testwith no flags, and gotAborting test run: test run timeout of 1 milliseconds exceeded/Test Run Aborted.Then restored it to 900000. That provesRunSettingsFilePathis honored on a flagless invocation, which is the whole claim.Data collector 'Blame' message: All tests finished running, Sequence file will not be generated.Both XML files validated as well-formed, after the toolchain caught me twice writing
--inside an XML comment (illegal), once in the csproj and once in the runsettings.Not in scope
The upstream bug itself. Evidence for a future dotnet/runtime report is preserved outside the session scratchpad at
~/Documents/dotnet-hang-evidence(sample, gzip, and a README explaining the signature); related to dotnet/runtime#66759 on Apple thread-suspension limits. This PR just makes the symptom self-terminating instead of open-ended.