Add CoreCLR GC pause duration histogram metrics - #133941
matyaskollert wants to merge 2 commits into
Conversation
Capture existing GC pause contributions in a bounded native queue and deliver generation- and type-tagged measurements through System.Runtime. Add overflow reporting, subscription-driven activation, and lifecycle and accounting tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Azure Pipelines: Successfully started running 4 pipeline(s). 12 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Tagging subscribers to this area: @steveisok, @dotnet/area-system-diagnostics-tracing |
|
|
Tagging subscribers to this area: @anicka-net, @dotnet/gc |
|
|
||
| internal static partial class RuntimeMetrics | ||
| { | ||
| private const string GCPauseReportingTypeName = "System.GCPauseReporting, System.Private.CoreLib"; |
There was a problem hiding this comment.
System.Diagnostics.DiagnosticSource is a nuget package independent on the runtime. It should depend on the runtime only through public APIs.
This is introducing de-facto public API without going through the proper process for public APIs.
There was a problem hiding this comment.
I will update the file once the public API is properly defined. Thank you
| uint64_t gc_heap::suspended_start_time = 0; | ||
| uint64_t gc_heap::end_gc_time = 0; | ||
| uint64_t gc_heap::total_suspended_time = 0; | ||
| #ifndef FEATURE_NATIVEAOT |
There was a problem hiding this comment.
FEATURE_NATIVEAOT ifdefs in the GC are anti-pattern. The GC should be identical for both NAOT and !NAOT.
There was a problem hiding this comment.
Does this mean that the new Histogram feature should be available in NativeAOT applications as well or just that the FEATURE_NATIVEAOT ifdefs should not be used in the GC?
There was a problem hiding this comment.
Does this mean that the new Histogram feature should be available in NativeAOT applications as well
I think so. What would be the rationale for excluding it for NAOT?
There was a problem hiding this comment.
I will include it for NAOT as well. Thank you
|
Looking through the implementation, I agree that it should publish this metric as a histogram as proposed, question is the "best" way to produce that histogram given the frequency and usage of updates and the sensitivity of the producer's location (GC during STW) and consumption of circular buffer from managed code (that could trigger new GC). This replicates a pattern we already have in runtime where native runtime code can safely emit events that could be consumed directly by managed code. Has an event-backed transport been considered as an alternative to the custom pause-record queue? Could the GC emit a self-contained event through Microsoft-Windows-DotNETRuntime, carrying the same duration and attribution fields under a dedicated keyword? An in-process EventListener in System.Diagnostics.DiagnosticSource could forward each record to the histogram. This would preserve the GC's accounting values without correlating existing GC start/suspend/restart events. The in-process EventListener path reads event instances directly from EventPipe session buffers; there is no .nettrace serialization/parsing round trip. There is still buffer management and managed payload decoding/allocation, but the custom transport also needs buffering, signaling, native-to-managed reads, and dispatch. What overhead budget are we targeting here? These measurements occur per GC pause contribution, rather than per allocation or object, so I would expect substantially lower event volume than many existing tracing scenarios. High-frequency Gen 0 collections still deserve measurement, and producer overhead during a pause matters, but it would be useful to establish whether this requires a specialized transport. Reusing EventPipe could avoid the dedicated GC queue, configure/drain/wait QCalls, reporting worker, and private CoreLib callback bridge. The adapter would live in DiagnosticSource, using EventListener, with no dependency from the GC on DiagnosticSource or metrics APIs. A self-contained pause event would also be useful beyond this histogram—for example, for tracing individual pauses and correlating them with application activity. That would expose the GC’s authoritative pause measurements through the existing diagnostics infrastructure, with the histogram as one consumer rather than the sole purpose of a dedicated transport. I would lean toward evaluating that approach first, including producer overhead, managed allocations, delivery latency/loss, and NativeAOT compatibility with and without tracing support. If it cannot meet those requirements, that would provide a concrete justification for the custom transport. |
|
@lateralusX thank you! The approach you've suggested sounds less intrusive to me. |
|
I tried the EventPipe approach locally, using a self-contained runtime event under a dedicated keyword and an EventListener in DiagnosticSource to forward it to the histogram. I can share the implementation in a separate draft PR for comparison. Most allocation-driven results were close. In the Windows x64 GCPerfSim runs, throughput was about 1.2% lower with custom and 2.1% lower with EventPipe compared with the unmodified runtime. The Server request-style profile did not show a clear throughput change. EventPipe added managed allocations, but the increase was small relative to the workload's allocations. The main differences I observed were:
I have the detailed tables and methodology written up separately. I didn't have a specific overhead or loss budget in mind when starting this, so I'd appreciate your guidance on which of these trade-offs matter most here and whether the EventPipe approach seems worth pursuing. |
|
@matyaskollert yes please, make a separate draft PR for EventPipe-based version. I agree that may be better approach. |
Keep the GC design document focused on its high-level architecture. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
I have opened a new PR #134121 with an alternative implementation using EventPipe as suggested above. |
dotnet.gc.pause.timereports a cumulative total, so it cannot distinguish many short pauses from a few long ones. This adds a histogram of GC-accounted pause contributions with generation and collection-type attributes.Related to #125753.
Approach
dotnet.gc.pause.duration(Histogram<double>, seconds), tagged withgc.heap.generation(gen0,gen1,gen2) andgc.pause.type(blocking,background).dotnet.gc.pause.droppedto report buffer overflow instead of blocking GC. Keepdotnet.gc.pause.timeunchanged.Performance
Compared periodic and event-driven batching against a preserved Release baseline on Windows x64. Event-driven delivery was selected because it avoided periodic idle activity and performed better in the allocation-driven Workstation workload.
The final high-rate forced-Gen-0 runs observed overflow, including 6,052 dropped records with the aggregating listener. Allocation-driven and background process trials had no drops, and recorded-duration sums matched native cumulative pause deltas. These are workload-specific measurements, not zero-overhead or lossless-delivery claims. A subsequent code-size simplification was compared against the prior implementation with 24 BenchmarkDotNet cases; ratios were 0.97-1.03, within observed variation.
Validation
GetTotalPauseDurationandGetGCMemoryInforuntime tests passed in Workstation and Server modes.Review notes
This is a draft for implementation and metric-schema review. Samples preserve the GC's existing per-collection accounting: collections sharing a suspension remain separately attributed, and background pause contributions remain separate samples. Delivery is asynchronous and bounded. Callbacks run on a background thread and must not throw; unhandled callback exceptions retain normal process-fatal behavior. The histogram schema, companion overflow counter, and these delivery semantics need agreement before merging.