Skip to content

Watchdog the macOS test-host hang from outside the process - #427

Merged
erikdarlingdata merged 1 commit into
devfrom
test-hang-watchdog
Aug 11, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
test-hang-watchdog

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

What does this PR do?

Stops PlanViewer.Core.Tests from spinning a full core indefinitely when it wedges on macOS ARM64.

It happened twice unprompted today: three test hosts at ~108% CPU each, one wedged 22 minutes and still going when sampled. It was noticed because the laptop got hot, which is not a monitoring strategy.

Why nothing in the test host can fix this

This is a CoreCLR GC-suspension livelock (.NET 10.0.8, macOS ARM64), not a slow or hanging test. From a 1 ms sample of a wedged host:

  • GCToEEInterface::SuspendEE → ThreadSuspend::SuspendEE in all 2451 samples — the runtime is trying to suspend for GC and never completes.
  • The victim thread was interrupted where CheckActivationSafePoint can never succeed, so the runtime re-sends activations forever.
  • The mach-exception thread churns thread_get_state / thread_set_state, which is what burns the core.

So xUnit timeouts and in-test CancelAfter cannot fire — the execution engine itself is suspended; that is the livelock. dotnet-stack and EventPipe hang for the same reason. Only an out-of-process watchdog works, and both mechanisms below are enforced by vstest.console, which is a separate process from the host it watches.

Two layers, and the second is the one that matters

  1. The three CI invocations (ci, nightly, release) pass the blame-hang options, which produce a sequence file naming the test that wedged.
  2. RunSettingsFilePath in the test csproj points at hang-watchdog.runsettings, giving a bare dotnet test a TestSessionTimeout backstop with no flags to remember.

Layer 2 is the important one: both incidents were plain local runs. A fix living only in the workflow files would protect the machine that was never at risk and leave the laptop unprotected — which is precisely how this recurred after being root-caused the first time.

15 minutes is far above any real runtime here (individual tests are milliseconds; the suite is normally well under a minute), so it only fires on a genuine wedge. It's a backstop, not a performance budget.

How was this tested?

A silently-ignored runsettings looks exactly like a working one, so I verified the mechanism rather than the file's existence:

  • Temporarily set TestSessionTimeout to 1, ran a bare dotnet test with no flags, and got Aborting test run: test run timeout of 1 milliseconds exceeded / Test Run Aborted. Then restored it to 900000. That proves RunSettingsFilePath is honored on a flagless invocation, which is the whole claim.
  • On a normal run the blame collector reports itself active: Data collector 'Blame' message: All tests finished running, Sequence file will not be generated.
  • Tests pass (2 passed, 1 skipped — the skip is a Windows-only pin), and zero test hosts survived any run.

Both XML files validated as well-formed, after the toolchain caught me twice writing -- inside an XML comment (illegal), once in the csproj and once in the runsettings.

Not in scope

The upstream bug itself. Evidence for a future dotnet/runtime report is preserved outside the session scratchpad at ~/Documents/dotnet-hang-evidence (sample, gzip, and a README explaining the signature); related to dotnet/runtime#66759 on Apple thread-suspension limits. This PR just makes the symptom self-terminating instead of open-ended.

PlanViewer.Core.Tests can wedge in a CoreCLR GC-suspension livelock on macOS ARM64
(.NET 10.0.8) and spin a full core until someone notices the heat. It happened
twice unprompted today: three hosts at ~108% CPU each, one of them wedged 22
minutes and still going when sampled. Root-caused from a sample: SuspendEE present
in all 2451 samples, the victim thread interrupted where CheckActivationSafePoint
can never succeed, and the mach-exception thread churning thread_get_state /
thread_set_state forever.

The load-bearing fact is that nothing inside the test host can stop it. xUnit
timeouts and in-test CancelAfter cannot fire because the execution engine itself is
suspended - that IS the livelock - and dotnet-stack and EventPipe hang for the same
reason. Only an out-of-process watchdog works, and both mechanisms here are
enforced by vstest.console, a separate process from the host it watches.

Two layers, deliberately:

- The three CI invocations (ci, nightly, release) now pass the blame-hang options,
  which is what produces a sequence file naming the test that wedged.
- RunSettingsFilePath in the test csproj points at hang-watchdog.runsettings, so a
  bare `dotnet test` with no flags gets a TestSessionTimeout backstop. This is the
  layer that matters most: BOTH incidents were plain local runs, so a fix living
  only in the workflow files would protect the machine that was never at risk and
  leave the laptop unprotected.

15 minutes is far above any real runtime (individual tests are milliseconds, the
whole suite normally well under a minute), so it only fires on a genuine wedge.

Verified rather than assumed, because a silently-ignored runsettings would look
exactly like a working one: temporarily set TestSessionTimeout to 1ms and a bare
`dotnet test` reported "Aborting test run: test run timeout of 1 milliseconds
exceeded", then restored it. The blame collector confirms itself active on a normal
run ("All tests finished running, Sequence file will not be generated"). Zero test
hosts survived any run.

Two XML-comment errors on the way in, both caught by the toolchain rather than by
me: '--' is illegal inside an XML comment, and I had written the flag names and a
dash-separated aside into comments in both the csproj and the runsettings.

Evidence for a future dotnet/runtime report is preserved outside the session
scratchpad at ~/Documents/dotnet-hang-evidence (sample, gzip, and a README
explaining the signature); related to dotnet/runtime#66759.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Merging on build-and-test and check-branches, and I want to be explicit that the green review check on this PR is not evidence of a review.

While checking it I found the review gate has been skipping repo-wide since 2026-08-03 and reporting success while doing it — filed as #428. Root cause is mine: #423 merged the Dependabot-skip into dev's claude-code-review.yml, main still has the 2026-07-26 version, and the action validates the running workflow against the default branch. Every PR targets dev, so validation fails and the action exits conclusion=success without reading anything. The tell was duration: 12 seconds, against ~4-5 minutes for a real review of comparable size in the sibling repo.

This PR would have skipped review regardless, since it modifies three workflow files and that always fails the identity check by design.

What I'm relying on instead: build-and-test genuinely ran the suite and passed, and the watchdog mechanism itself is verified empirically rather than by inspection — a bare dotnet test with TestSessionTimeout temporarily set to 1 reported Aborting test run: test run timeout of 1 milliseconds exceeded, proving RunSettingsFilePath is honored on a flagless invocation. That's the claim this PR makes.

Merging because it stops a recurring problem that has burned three cores twice today and is currently detected by noticing the laptop is hot.

@erikdarlingdata
erikdarlingdata merged commit 897d241 into dev Aug 11, 2026
3 checks passed
@erikdarlingdata
erikdarlingdata deleted the test-hang-watchdog branch August 21, 2026 10:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant