Skip to content

[RFC] Server-initiated session cancellation in the dotnet test IPC protocol #8691

Description

RFC NNN - Server-initiated session cancellation in the dotnet test IPC protocol

  • Approved in principle
  • Under discussion
  • Implementation
  • Shipped

Summary

Today the IPC protocol between dotnet test (server / SDK) and a Microsoft Testing Platform (MTP) test app (client) is strictly client→server request/response: the test app pushes events (handshake, test results, file artifacts, session events) and the SDK replies with an empty VoidResponse. The SDK has no way to ask the test app to stop.

This RFC proposes adding a server-initiated cancel session primitive so that dotnet test-level features can cooperatively terminate test apps mid-run.

Motivation

Several dotnet test UX features need to be applied across all test apps in a solution / list of projects, not per-app:

Without a protocol primitive, dotnet test can only:

  1. Let the apps run to completion (today's behavior — wastes time, misleading semantics for "max failed tests"), or
  2. SIGKILL them (drops reporting, breaks trx, no chance for cleanup hooks).

What's missing is a third option: ask politely.

Detailed design

Capability discovery (handshake)

Add a new optional handshake property name in Constants.HandshakeMessagePropertyNames:

ServerControlChannel = "10"   // value: "True" / "False"

When both sides advertise ServerControlChannel=True in the handshake message, the feature is enabled for that connection. This is orthogonal to ProtocolVersion — capability-based negotiation avoids the well-known SemVer-string-compare pitfalls (e.g., "1.10.0" < "1.2.0" when compared as strings).

ProtocolConstants.Version should still bump to 1.1.0 for the version-aware path described in "Negotiation" below, but the cancel feature gate is the capability flag, not the version string.

Two implementation options

The crux of the RFC is how the server signal reaches the test app. Both options share the capability negotiation above; they differ only in the transport.

Option A — Ack-piggyback on VoidResponse

Introduce a new response type that replaces VoidResponse for clients that advertise the capability:

public sealed record ServerAckResponse(byte Flags) : IResponse;

[Flags]
internal static class ServerAckFlags
{
    public const byte None          = 0;
    public const byte CancelSession = 1 << 0;
    // reserved bits for future signals (drain, pause, ...)
}

Assign serializer id 11 (next free). The SDK chooses ServerAckResponse vs VoidResponse per-connection based on the negotiated capability so old clients keep getting empty replies.

When the SDK wants to cancel, it stashes a "cancel" intent and sets CancelSession=1 on the next outgoing ack. The test app, on receiving any ServerAckResponse with CancelSession=1, triggers ITestApplicationCancellationTokenSource.CancellationTokenSource.Cancel().

Pros: small wire change, single connection, no new pipe to manage, no new lifecycle.

Cons (the big one): purely opportunistic. If the test app is hung or has gone silent (e.g., a hanging adapter / a slow test that doesn't emit TestInProgress events), the SDK has no message to piggyback on. This makes Option A a poor fit for --timeout and a leaky fit for --maximum-failed-tests (we may have already reached N but the next-to-die test app may not send another message for minutes).

Option B — Reverse control pipe

The SDK creates a second NamedPipeServer and advertises its name to each test app in the handshake response under a new property:

ServerControlPipeName = "11"   // value: "<unique-pipe-name>"

The test app, on seeing this property, opens a NamedPipeClient to that pipe and parks a single long-lived RequestReplyAsync<WaitForServerControlRequest, ServerControlMessage> on it. The SDK completes that request whenever it wants to signal the app:

public sealed record ServerControlMessage(byte Kind);  // 1 = CancelSession, ...

When ServerControlMessage.Kind == CancelSession arrives, the test app cancels and the long-poll loop ends. (Connection close also implies "host went away → cancel".)

Pros: true server push, works for hung/silent test apps, naturally extensible to drain/pause/etc., dead-host detection comes "for free" via pipe disconnect.

Cons: more code (extra NamedPipeServer on the SDK side, extra NamedPipeClient and a long-poll background task on the test app side), more lifecycle to reason about (what if the control pipe drops mid-session?), one extra OS handle per connection.

Recommendation

Option B, optionally with Option A as an opportunistic optimization later (no second pipe needed for fast cancellation of busy test apps).

The motivating scenarios (--timeout, --maximum-failed-tests) all benefit from being able to cancel a stalled or silent test app. If we ship Option A first and then realize we need B anyway, we've added two protocol features for the work of one.

Negotiation: also fix the version handshake

Independent of cancel, the existing handshake has a latent bug worth fixing at the same time:

// IsCompatibleProtocolAsync (test-app side)
bool isCompatible = supportedProtocolVersions.Split(';').Contains(protocolVersion);

The returned value is discarded — there is no "the negotiated version is X" state on either side. With this RFC we should:

  1. Have the server return a single negotiated version in the handshake reply (the highest version present in both sides' preference lists), not a ;-separated set.
  2. Persist NegotiatedProtocolVersion and NegotiatedCapabilities per-connection on both sides.
  3. Gate every protocol decision (e.g. "should I send ServerAckResponse or VoidResponse?") on this per-connection state, never on the local ProtocolConstants.Version.

This is needed for clean forward-compat regardless of which cancel option we pick.

Behavior after cancel

The RFC should spec what happens after the test app observes cancel:

  • The test app stops scheduling new tests and triggers cooperative cancellation of in-flight tests via ITestApplicationCancellationTokenSource.
  • The test app still attempts to send TestSessionEnd, any final test results that complete (or are reported as canceled), and FileArtifactMessages on a separate non-canceled token before exiting. Without this, the SDK loses the trx / logs / coverage from anything that did complete.
  • The SDK retains the right to escalate to Process.Kill after a grace period (suggested 30s, configurable later). Cancellation is cooperative — adapters that don't honor the token won't honor this either.
  • Exit code: when the SDK cancels a test app via this primitive, the SDK treats the app's exit as "canceled by host", distinct from "test crashed". TBD whether we need a new MTP exit code for this or whether the SDK simply records the cause out-of-band.

Drawbacks

  • Two-way protocol is conceptually heavier than today's one-way push. Adding it means we need to maintain it carefully — every future protocol consumer (IDEs, third-party tools, alternative non-.NET test apps) needs to handle the new shape.
  • Mixed-version matrix grows: new-SDK ↔ old-test-app, old-SDK ↔ new-test-app, new ↔ new with cancel-on-Nth-message all need test coverage. The protocol files are also duplicated in dotnet/sdk, so PRs must land in lockstep.
  • For Option B specifically: one extra pipe per connection on Windows; on Linux/macOS the same with Unix domain sockets. Not free.

Alternatives

  • Status quo + Process.Kill. Works but drops reporting and breaks trx. Acceptable for Ctrl+C (already what we do), not acceptable as the default behavior for --maximum-failed-tests or --timeout.
  • Out-of-band signal (e.g., SIGINT / CTRL_BREAK_EVENT). Cross-platform reliability is poor and existing user Ctrl+C handlers in test code would also fire. Doesn't extend to drain/pause.
  • Push the per-test-app --maximum-failed-tests=N down before the run. Already exists per-app, but it cannot express "total across the solution = N" without the host coordinating across apps.

Compatibility

  • Wire: no breaking change for clients/servers that don't negotiate the new capability. Old test apps connecting to a new SDK continue to receive VoidResponse and never see a control pipe name. Old SDKs connecting to a new test app (e.g., a user upgrades MTP without upgrading their SDK) ignore the new handshake property → test app sees the SDK didn't advertise the capability → no cancel attempted.
  • Behavior: when both sides support cancel, this enables an SDK→test-app channel that didn't exist before. Adapters that don't honor the cancellation token will not see behavior change today (cooperative); we should call this out in docs.
  • Duplicated protocol files: any change here must land in both microsoft/testfx and dotnet/sdk in coordinated PRs.

Unresolved questions

  1. Option A vs B — which transport do we want? (RFC recommends B.)
  2. Non-.NET test app clients — are there any (e.g., a JS or native MTP host) that would have to implement the new shape? If yes, the answer leans further toward B (simpler to implement a one-shot long-poll than to retrofit ack semantics on every reply).
  3. Pipe-name exchange location — handshake response (proposed) vs a separate post-handshake "subscribe" request.
  4. Grace period before SIGKILL — fixed default vs SDK option vs per-test-app override.
  5. Exit-code semantics — do we need a new MTP exit code for "canceled by host" (distinct from user Ctrl+C and from crash)?
  6. Version negotiation fix — should this ship as a separate prereq RFC/PR, or roll it into this one?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/mtpMicrosoft.Testing.Platform core library.type/rfcRequest for comments.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions