Skip to content

Add CoreCLR GC pause duration histogram metrics - #133941

Draft
matyaskollert wants to merge 2 commits into
dotnet:mainfrom
matyaskollert:matyaskollert-gc-collect-phase
Draft

matyaskollert wants to merge 2 commits into
dotnet:mainfrom
matyaskollert:matyaskollert-gc-collect-phase

Conversation

@matyaskollert

Copy link
Copy Markdown
Member

dotnet.gc.pause.time reports a cumulative total, so it cannot distinguish many short pauses from a few long ones. This adds a histogram of GC-accounted pause contributions with generation and collection-type attributes.

Related to #125753.

Approach

  • Add dotnet.gc.pause.duration (Histogram<double>, seconds), tagged with gc.heap.generation (gen0, gen1, gen2) and gc.pause.type (blocking, background).
  • Capture the durations where CoreCLR already updates its pause accounting. A bounded native queue feeds an event-driven managed consumer, without additional clock reads or managed callbacks during collection. Capture follows histogram subscriptions and does not poll while idle.
  • Add dotnet.gc.pause.dropped to report buffer overflow instead of blocking GC. Keep dotnet.gc.pause.time unchanged.
  • Use an internal CoreLib bridge with standalone-GC version checks. No new public C# APIs; reporting is CoreCLR-only.

Performance

Compared periodic and event-driven batching against a preserved Release baseline on Windows x64. Event-driven delivery was selected because it avoided periodic idle activity and performed better in the allocation-driven Workstation workload.

Workload Observed impact versus the original runtime
Allocation-driven, Workstation GC About 0.9% slower across three rotating-order trials
Allocation-driven, four-heap Server GC Within run-to-run noise
Repeatedly forced small Gen 0 collections About 8-14% slower with a histogram listener

The final high-rate forced-Gen-0 runs observed overflow, including 6,052 dropped records with the aggregating listener. Allocation-driven and background process trials had no drops, and recorded-duration sums matched native cumulative pause deltas. These are workload-specific measurements, not zero-overhead or lossless-delivery claims. A subsequent code-size simplification was compared against the prior implementation with 24 BenchmarkDotNet cases; ratios were 0.97-1.03, within observed variation.

Validation

  • Checked and Release CoreCLR builds; DiagnosticSource multi-target build.
  • 459 DiagnosticSource test cases passed on Checked CoreCLR after simplification.
  • Existing GetTotalPauseDuration and GetGCMemoryInfo runtime tests passed in Workstation and Server modes.
  • NativeAOT compilation compatibility and new-VM/old-standalone-GC capability checks.
  • Coverage for native accounting, background contributions, listener lifecycle, overflow, registration backlog, callback-triggered GC/disposal, and rapid resubscription.

Review notes

This is a draft for implementation and metric-schema review. Samples preserve the GC's existing per-collection accounting: collections sharing a suspension remain separately attributed, and background pause contributions remain separate samples. Delivery is asynchronous and bounded. Callbacks run on a background thread and must not throw; unhandled callback exceptions retain normal process-fatal behavior. The histogram schema, companion overflow counter, and these delivery semantics need agreement before merging.

Capture existing GC pause contributions in a bounded native queue and deliver generation- and type-tagged measurements through System.Runtime. Add overflow reporting, subscription-driven activation, and lifecycle and accounting tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 4 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @steveisok, @dotnet/area-system-diagnostics-tracing
See info in area-owners.md if you want to be subscribed.

@dotnet-policy-service dotnet-policy-service Bot added the linkable-framework Issues associated with delivering a linker friendly framework label Sep 15, 2026
@matyaskollert

Copy link
Copy Markdown
Member Author
@dotnet-policy-service agree company="Microsoft"

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @anicka-net, @dotnet/gc
See info in area-owners.md if you want to be subscribed.


internal static partial class RuntimeMetrics
{
private const string GCPauseReportingTypeName = "System.GCPauseReporting, System.Private.CoreLib";

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

System.Diagnostics.DiagnosticSource is a nuget package independent on the runtime. It should depend on the runtime only through public APIs.

This is introducing de-facto public API without going through the proper process for public APIs.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will update the file once the public API is properly defined. Thank you

Comment thread docs/design/coreclr/botr/garbage-collection.md Outdated
Comment thread src/coreclr/gc/gc.cpp
uint64_t gc_heap::suspended_start_time = 0;
uint64_t gc_heap::end_gc_time = 0;
uint64_t gc_heap::total_suspended_time = 0;
#ifndef FEATURE_NATIVEAOT

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FEATURE_NATIVEAOT ifdefs in the GC are anti-pattern. The GC should be identical for both NAOT and !NAOT.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this mean that the new Histogram feature should be available in NativeAOT applications as well or just that the FEATURE_NATIVEAOT ifdefs should not be used in the GC?

@jkotas jkotas Sep 15, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this mean that the new Histogram feature should be available in NativeAOT applications as well

I think so. What would be the rationale for excluding it for NAOT?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will include it for NAOT as well. Thank you

@lateralusX

lateralusX commented Sep 16, 2026 •

Copy link
Copy Markdown
Member

Looking through the implementation, I agree that it should publish this metric as a histogram as proposed, question is the "best" way to produce that histogram given the frequency and usage of updates and the sensitivity of the producer's location (GC during STW) and consumption of circular buffer from managed code (that could trigger new GC). This replicates a pattern we already have in runtime where native runtime code can safely emit events that could be consumed directly by managed code. Has an event-backed transport been considered as an alternative to the custom pause-record queue?

Could the GC emit a self-contained event through Microsoft-Windows-DotNETRuntime, carrying the same duration and attribution fields under a dedicated keyword? An in-process EventListener in System.Diagnostics.DiagnosticSource could forward each record to the histogram. This would preserve the GC's accounting values without correlating existing GC start/suspend/restart events.

The in-process EventListener path reads event instances directly from EventPipe session buffers; there is no .nettrace serialization/parsing round trip. There is still buffer management and managed payload decoding/allocation, but the custom transport also needs buffering, signaling, native-to-managed reads, and dispatch.

What overhead budget are we targeting here? These measurements occur per GC pause contribution, rather than per allocation or object, so I would expect substantially lower event volume than many existing tracing scenarios. High-frequency Gen 0 collections still deserve measurement, and producer overhead during a pause matters, but it would be useful to establish whether this requires a specialized transport.

Reusing EventPipe could avoid the dedicated GC queue, configure/drain/wait QCalls, reporting worker, and private CoreLib callback bridge. The adapter would live in DiagnosticSource, using EventListener, with no dependency from the GC on DiagnosticSource or metrics APIs.

A self-contained pause event would also be useful beyond this histogram—for example, for tracing individual pauses and correlating them with application activity. That would expose the GC’s authoritative pause measurements through the existing diagnostics infrastructure, with the histogram as one consumer rather than the sole purpose of a dedicated transport.

I would lean toward evaluating that approach first, including producer overhead, managed allocations, delivery latency/loss, and NativeAOT compatibility with and without tracing support. If it cannot meet those requirements, that would provide a concrete justification for the custom transport.

@janvorli

Copy link
Copy Markdown
Member

@lateralusX thank you! The approach you've suggested sounds less intrusive to me.

@matyaskollert

Copy link
Copy Markdown
Member Author

I tried the EventPipe approach locally, using a self-contained runtime event under a dedicated keyword and an EventListener in DiagnosticSource to forward it to the histogram. I can share the implementation in a separate draft PR for comparison.

Most allocation-driven results were close. In the Windows x64 GCPerfSim runs, throughput was about 1.2% lower with custom and 2.1% lower with EventPipe compared with the unmodified runtime. The Server request-style profile did not show a clear throughput change. EventPipe added managed allocations, but the increase was small relative to the workload's allocations.

The main differences I observed were:

  • Producer cost: in synthetic tests with spaced collections, writing a record averaged roughly 4-7 us with custom versus 21-34 us with EventPipe.
  • Delivery latency: in those paced tests, custom typically delivered records in tens of microseconds, while EventPipe's median was around 15 ms.
  • Buffering and loss: EventPipe retained more records during large bursts, though its buffer budget is much larger. Custom exposes a dropped-record count; the current EventPipe path does not. In a deliberately blocked-reader test, enabling and disabling another listener to the native runtime provider discarded queued EventPipe records, while the custom queue retained them.

I have the detailed tables and methodology written up separately. I didn't have a specific overhead or loss budget in mind when starting this, so I'd appreciate your guidance on which of these trade-offs matter most here and whether the EventPipe approach seems worth pursuing.

@kkokosa

kkokosa commented Sep 16, 2026

Copy link
Copy Markdown
Member

@matyaskollert yes please, make a separate draft PR for EventPipe-based version. I agree that may be better approach.

Keep the GC design document focused on its high-level architecture.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@matyaskollert

Copy link
Copy Markdown
Member Author

I have opened a new PR #134121 with an alternative implementation using EventPipe as suggested above.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-GC-coreclr linkable-framework Issues associated with delivering a linker friendly framework

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants