RFC NNN - Server-initiated session cancellation in the dotnet test IPC protocol
Summary
Today the IPC protocol between dotnet test (server / SDK) and a Microsoft Testing Platform (MTP) test app (client) is strictly client→server request/response: the test app pushes events (handshake, test results, file artifacts, session events) and the SDK replies with an empty VoidResponse. The SDK has no way to ask the test app to stop.
This RFC proposes adding a server-initiated cancel session primitive so that dotnet test-level features can cooperatively terminate test apps mid-run.
Motivation
Several dotnet test UX features need to be applied across all test apps in a solution / list of projects, not per-app:
Without a protocol primitive, dotnet test can only:
- Let the apps run to completion (today's behavior — wastes time, misleading semantics for "max failed tests"), or
- SIGKILL them (drops reporting, breaks trx, no chance for cleanup hooks).
What's missing is a third option: ask politely.
Detailed design
Capability discovery (handshake)
Add a new optional handshake property name in Constants.HandshakeMessagePropertyNames:
ServerControlChannel = "10" // value: "True" / "False"
When both sides advertise ServerControlChannel=True in the handshake message, the feature is enabled for that connection. This is orthogonal to ProtocolVersion — capability-based negotiation avoids the well-known SemVer-string-compare pitfalls (e.g., "1.10.0" < "1.2.0" when compared as strings).
ProtocolConstants.Version should still bump to 1.1.0 for the version-aware path described in "Negotiation" below, but the cancel feature gate is the capability flag, not the version string.
Two implementation options
The crux of the RFC is how the server signal reaches the test app. Both options share the capability negotiation above; they differ only in the transport.
Option A — Ack-piggyback on VoidResponse
Introduce a new response type that replaces VoidResponse for clients that advertise the capability:
public sealed record ServerAckResponse(byte Flags) : IResponse;
[Flags]
internal static class ServerAckFlags
{
public const byte None = 0;
public const byte CancelSession = 1 << 0;
// reserved bits for future signals (drain, pause, ...)
}
Assign serializer id 11 (next free). The SDK chooses ServerAckResponse vs VoidResponse per-connection based on the negotiated capability so old clients keep getting empty replies.
When the SDK wants to cancel, it stashes a "cancel" intent and sets CancelSession=1 on the next outgoing ack. The test app, on receiving any ServerAckResponse with CancelSession=1, triggers ITestApplicationCancellationTokenSource.CancellationTokenSource.Cancel().
Pros: small wire change, single connection, no new pipe to manage, no new lifecycle.
Cons (the big one): purely opportunistic. If the test app is hung or has gone silent (e.g., a hanging adapter / a slow test that doesn't emit TestInProgress events), the SDK has no message to piggyback on. This makes Option A a poor fit for --timeout and a leaky fit for --maximum-failed-tests (we may have already reached N but the next-to-die test app may not send another message for minutes).
Option B — Reverse control pipe
The SDK creates a second NamedPipeServer and advertises its name to each test app in the handshake response under a new property:
ServerControlPipeName = "11" // value: "<unique-pipe-name>"
The test app, on seeing this property, opens a NamedPipeClient to that pipe and parks a single long-lived RequestReplyAsync<WaitForServerControlRequest, ServerControlMessage> on it. The SDK completes that request whenever it wants to signal the app:
public sealed record ServerControlMessage(byte Kind); // 1 = CancelSession, ...
When ServerControlMessage.Kind == CancelSession arrives, the test app cancels and the long-poll loop ends. (Connection close also implies "host went away → cancel".)
Pros: true server push, works for hung/silent test apps, naturally extensible to drain/pause/etc., dead-host detection comes "for free" via pipe disconnect.
Cons: more code (extra NamedPipeServer on the SDK side, extra NamedPipeClient and a long-poll background task on the test app side), more lifecycle to reason about (what if the control pipe drops mid-session?), one extra OS handle per connection.
Recommendation
Option B, optionally with Option A as an opportunistic optimization later (no second pipe needed for fast cancellation of busy test apps).
The motivating scenarios (--timeout, --maximum-failed-tests) all benefit from being able to cancel a stalled or silent test app. If we ship Option A first and then realize we need B anyway, we've added two protocol features for the work of one.
Negotiation: also fix the version handshake
Independent of cancel, the existing handshake has a latent bug worth fixing at the same time:
// IsCompatibleProtocolAsync (test-app side)
bool isCompatible = supportedProtocolVersions.Split(';').Contains(protocolVersion);
The returned value is discarded — there is no "the negotiated version is X" state on either side. With this RFC we should:
- Have the server return a single negotiated version in the handshake reply (the highest version present in both sides' preference lists), not a
;-separated set.
- Persist
NegotiatedProtocolVersion and NegotiatedCapabilities per-connection on both sides.
- Gate every protocol decision (e.g. "should I send
ServerAckResponse or VoidResponse?") on this per-connection state, never on the local ProtocolConstants.Version.
This is needed for clean forward-compat regardless of which cancel option we pick.
Behavior after cancel
The RFC should spec what happens after the test app observes cancel:
- The test app stops scheduling new tests and triggers cooperative cancellation of in-flight tests via
ITestApplicationCancellationTokenSource.
- The test app still attempts to send
TestSessionEnd, any final test results that complete (or are reported as canceled), and FileArtifactMessages on a separate non-canceled token before exiting. Without this, the SDK loses the trx / logs / coverage from anything that did complete.
- The SDK retains the right to escalate to
Process.Kill after a grace period (suggested 30s, configurable later). Cancellation is cooperative — adapters that don't honor the token won't honor this either.
- Exit code: when the SDK cancels a test app via this primitive, the SDK treats the app's exit as "canceled by host", distinct from "test crashed". TBD whether we need a new MTP exit code for this or whether the SDK simply records the cause out-of-band.
Drawbacks
- Two-way protocol is conceptually heavier than today's one-way push. Adding it means we need to maintain it carefully — every future protocol consumer (IDEs, third-party tools, alternative non-.NET test apps) needs to handle the new shape.
- Mixed-version matrix grows: new-SDK ↔ old-test-app, old-SDK ↔ new-test-app, new ↔ new with cancel-on-Nth-message all need test coverage. The protocol files are also duplicated in dotnet/sdk, so PRs must land in lockstep.
- For Option B specifically: one extra pipe per connection on Windows; on Linux/macOS the same with Unix domain sockets. Not free.
Alternatives
- Status quo +
Process.Kill. Works but drops reporting and breaks trx. Acceptable for Ctrl+C (already what we do), not acceptable as the default behavior for --maximum-failed-tests or --timeout.
- Out-of-band signal (e.g.,
SIGINT / CTRL_BREAK_EVENT). Cross-platform reliability is poor and existing user Ctrl+C handlers in test code would also fire. Doesn't extend to drain/pause.
- Push the per-test-app
--maximum-failed-tests=N down before the run. Already exists per-app, but it cannot express "total across the solution = N" without the host coordinating across apps.
Compatibility
- Wire: no breaking change for clients/servers that don't negotiate the new capability. Old test apps connecting to a new SDK continue to receive
VoidResponse and never see a control pipe name. Old SDKs connecting to a new test app (e.g., a user upgrades MTP without upgrading their SDK) ignore the new handshake property → test app sees the SDK didn't advertise the capability → no cancel attempted.
- Behavior: when both sides support cancel, this enables an SDK→test-app channel that didn't exist before. Adapters that don't honor the cancellation token will not see behavior change today (cooperative); we should call this out in docs.
- Duplicated protocol files: any change here must land in both microsoft/testfx and dotnet/sdk in coordinated PRs.
Unresolved questions
- Option A vs B — which transport do we want? (RFC recommends B.)
- Non-.NET test app clients — are there any (e.g., a JS or native MTP host) that would have to implement the new shape? If yes, the answer leans further toward B (simpler to implement a one-shot long-poll than to retrofit ack semantics on every reply).
- Pipe-name exchange location — handshake response (proposed) vs a separate post-handshake "subscribe" request.
- Grace period before SIGKILL — fixed default vs SDK option vs per-test-app override.
- Exit-code semantics — do we need a new MTP exit code for "canceled by host" (distinct from user
Ctrl+C and from crash)?
- Version negotiation fix — should this ship as a separate prereq RFC/PR, or roll it into this one?
RFC NNN - Server-initiated session cancellation in the
dotnet testIPC protocolSummary
Today the IPC protocol between
dotnet test(server / SDK) and a Microsoft Testing Platform (MTP) test app (client) is strictly client→server request/response: the test app pushes events (handshake, test results, file artifacts, session events) and the SDK replies with an emptyVoidResponse. The SDK has no way to ask the test app to stop.This RFC proposes adding a server-initiated cancel session primitive so that
dotnet test-level features can cooperatively terminate test apps mid-run.Motivation
Several
dotnet testUX features need to be applied across all test apps in a solution / list of projects, not per-app:--maximum-failed-tests N—dotnet testaggregates failures across apps; once N is reached, the remaining (and currently running) apps should stop ASAP rather than continue running to completion. See dotnet test for MTP handling of "global" command-line options dotnet/sdk#49709 and the upcoming SDK PR for the post-hoc / exit-code-only first step.--timeout— when the wall-clock budget for the wholedotnet testinvocation is exceeded, currently the only lever isProcess.Kill, which produces no trx, no logs, no artifacts.Without a protocol primitive,
dotnet testcan only:What's missing is a third option: ask politely.
Detailed design
Capability discovery (handshake)
Add a new optional handshake property name in
Constants.HandshakeMessagePropertyNames:When both sides advertise
ServerControlChannel=Truein the handshake message, the feature is enabled for that connection. This is orthogonal toProtocolVersion— capability-based negotiation avoids the well-known SemVer-string-compare pitfalls (e.g.,"1.10.0" < "1.2.0"when compared as strings).ProtocolConstants.Versionshould still bump to1.1.0for the version-aware path described in "Negotiation" below, but the cancel feature gate is the capability flag, not the version string.Two implementation options
The crux of the RFC is how the server signal reaches the test app. Both options share the capability negotiation above; they differ only in the transport.
Option A — Ack-piggyback on
VoidResponseIntroduce a new response type that replaces
VoidResponsefor clients that advertise the capability:Assign serializer id 11 (next free). The SDK chooses
ServerAckResponsevsVoidResponseper-connection based on the negotiated capability so old clients keep getting empty replies.When the SDK wants to cancel, it stashes a "cancel" intent and sets
CancelSession=1on the next outgoing ack. The test app, on receiving anyServerAckResponsewithCancelSession=1, triggersITestApplicationCancellationTokenSource.CancellationTokenSource.Cancel().Pros: small wire change, single connection, no new pipe to manage, no new lifecycle.
Cons (the big one): purely opportunistic. If the test app is hung or has gone silent (e.g., a hanging adapter / a slow test that doesn't emit
TestInProgressevents), the SDK has no message to piggyback on. This makes Option A a poor fit for--timeoutand a leaky fit for--maximum-failed-tests(we may have already reached N but the next-to-die test app may not send another message for minutes).Option B — Reverse control pipe
The SDK creates a second
NamedPipeServerand advertises its name to each test app in the handshake response under a new property:The test app, on seeing this property, opens a
NamedPipeClientto that pipe and parks a single long-livedRequestReplyAsync<WaitForServerControlRequest, ServerControlMessage>on it. The SDK completes that request whenever it wants to signal the app:When
ServerControlMessage.Kind == CancelSessionarrives, the test app cancels and the long-poll loop ends. (Connection close also implies "host went away → cancel".)Pros: true server push, works for hung/silent test apps, naturally extensible to drain/pause/etc., dead-host detection comes "for free" via pipe disconnect.
Cons: more code (extra
NamedPipeServeron the SDK side, extraNamedPipeClientand a long-poll background task on the test app side), more lifecycle to reason about (what if the control pipe drops mid-session?), one extra OS handle per connection.Recommendation
Option B, optionally with Option A as an opportunistic optimization later (no second pipe needed for fast cancellation of busy test apps).
The motivating scenarios (
--timeout,--maximum-failed-tests) all benefit from being able to cancel a stalled or silent test app. If we ship Option A first and then realize we need B anyway, we've added two protocol features for the work of one.Negotiation: also fix the version handshake
Independent of cancel, the existing handshake has a latent bug worth fixing at the same time:
The returned value is discarded — there is no "the negotiated version is X" state on either side. With this RFC we should:
;-separated set.NegotiatedProtocolVersionandNegotiatedCapabilitiesper-connection on both sides.ServerAckResponseorVoidResponse?") on this per-connection state, never on the localProtocolConstants.Version.This is needed for clean forward-compat regardless of which cancel option we pick.
Behavior after cancel
The RFC should spec what happens after the test app observes cancel:
ITestApplicationCancellationTokenSource.TestSessionEnd, any final test results that complete (or are reported as canceled), andFileArtifactMessageson a separate non-canceled token before exiting. Without this, the SDK loses the trx / logs / coverage from anything that did complete.Process.Killafter a grace period (suggested 30s, configurable later). Cancellation is cooperative — adapters that don't honor the token won't honor this either.Drawbacks
Alternatives
Process.Kill. Works but drops reporting and breaks trx. Acceptable forCtrl+C(already what we do), not acceptable as the default behavior for--maximum-failed-testsor--timeout.SIGINT/CTRL_BREAK_EVENT). Cross-platform reliability is poor and existing userCtrl+Chandlers in test code would also fire. Doesn't extend to drain/pause.--maximum-failed-tests=Ndown before the run. Already exists per-app, but it cannot express "total across the solution = N" without the host coordinating across apps.Compatibility
VoidResponseand never see a control pipe name. Old SDKs connecting to a new test app (e.g., a user upgrades MTP without upgrading their SDK) ignore the new handshake property → test app sees the SDK didn't advertise the capability → no cancel attempted.Unresolved questions
Ctrl+Cand from crash)?