Skip to content

Use controller-backed TRX recovery by default #10792

Description

Summary

Run TRX in controller-backed mode whenever the platform supports a test-host controller. Do not introduce a TRX mode option initially: --report-trx should consistently select the reliable controller-backed implementation on supported platforms.

This does not move normal per-test TRX generation into the controller. The test host continues streaming and generating the report; the surviving controller owns recovery when the child exits before finalization.

Background and Motivation

TRX currently chooses its controller-coordinated recovery path by checking whether --crashdump is set. That creates several gaps:

  • --hangdump --report-trx already launches a child test host, but TRX does not join the controller recovery path.
  • --timeout --report-trx has no survivor unless another extension happens to force a controller.
  • a plain --report-trx run loses the final report if the test-host process crashes, even though completed results were flushed to the binary sidecar.
  • future controller extensions do not automatically provide TRX recovery.

The closures of #5145 and #6778 assumed that #8263 covered timeout and HangDump, but the current CrashDumpCommandLineOptions gate prevents those general cases from using TRX recovery.

PR #10799 implemented an earlier design as a stacked PR, but it was accidentally merged into the still-open #10797 branch rather than main. Its changes were subsequently reverted from #10797 to keep the timeout fix focused. This issue tracks the replacement implementation reviewed in #10797 (comment).

Proposed Feature

  1. When --report-trx is enabled on a platform and application model that supports a test-host controller, let TRX require the controller and enable its recovery components.
  2. Replace the direct --crashdump check in TrxModeHelpers with effective controller-presence state supplied by MTP.
  3. In the child process, use actual controller-presence state rather than knowledge of which extension caused isolation.
  4. On platforms where process restart is unavailable, such as browser/WASI, retain the existing in-process implementation as a compatibility fallback.
  5. Do not add --trx-mode or a user-selectable in-process opt-out initially. If concrete compatibility or performance evidence later requires an opt-out, design it separately.
  6. Document the supported-platform behavior and reliability tradeoff without exposing an unnecessary mode-selection surface.

The timeout lifecycle work in #10791 is a prerequisite. Controller-backed TRX must finalize under a separate bounded cleanup token and emit a failed report after timeout.

Acceptance coverage should include:

  • normal successful TRX generation through the controller-backed default;
  • test-host crash with only --report-trx;
  • HangDump-triggered termination;
  • timeout-triggered termination after some tests completed;
  • failed ResultSummary and run diagnostics for abnormal termination;
  • browser/WASI and unsupported application-model fallback behavior;
  • startup/performance measurements to quantify the additional process cost;
  • debugger, custom launcher, Native AOT, single-file, and packaged-app compatibility where applicable.

Existing output assertions must reflect the controller-backed artifact location. Tests should explicitly verify supported-platform controller-backed behavior rather than only validating defaults indirectly.

Alternative Designs

  • Only join an existing controller. Register TRX as a passive controller extension that activates when another extension already requires process restart. This fixes HangDump and similar combinations but leaves plain TRX crashes and --timeout --report-trx without a survivor.
  • Keep the current --crashdump gate. This preserves performance but makes recovery depend on an unrelated option and leaves known correctness gaps.
  • Add --trx-mode in-process|out-of-process. This was prototyped in Default TRX to controller-backed recovery #10799 but rejected as premature API surface. Start with one reliable behavior and add an opt-out only if evidence demonstrates a need.
  • Publish periodic atomic TRX snapshots from the in-process journal. This can survive whole-process-tree termination and avoids an extra process, but requires additional streaming/snapshot work. The Prototype crash-resilient TRX writes #10106 prototype showed journal plus atomic snapshots is the safe direction; it is complementary long-term work.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/mtpMicrosoft.Testing.Platform core library.area/mtp-extensionsMTP extensions (TrxReport, Retry, HtmlReport, ...).area/trxTRX report extension.type/breaking-changeBehavioral or API breaking change.

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions