Skip to content

JIT: Fix value profiling in optimized instrumented tiers - #134160

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-profile-values-r2r
Sep 18, 2026
Merged

EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-profile-values-r2r

Conversation

@EgorBo

@EgorBo EgorBo commented Sep 17, 2026 •

Copy link
Copy Markdown
Member

Collect Memmove/SequenceEqual length profiles in optimized instrumented tiers, including inline candidates.
This enables the existing PGO-driven unrolling with default tiering and ReadyToRun enabled.

using System.Runtime.CompilerServices;

for (int i = 0; i < 300; i++)
{
    CopyTo(new Span<char>(new char[40]), "Hello World");
    Thread.Sleep(16);
}

[MethodImpl(MethodImplOptions.NoInlining)]
void CopyTo(Span<char> dst, ReadOnlySpan<char> src) => src.CopyTo(dst);

Before (Windows x64, Tier1):

G_M28613_IG01:
       sub      rsp, 40
G_M28613_IG02:
       mov      rax, bword ptr [rdx]
       mov      r8d, dword ptr [rdx+0x08]
       mov      rdx, bword ptr [rcx]
       mov      ecx, dword ptr [rcx+0x08]
       cmp      r8d, ecx
       jg       SHORT G_M28613_IG04
       add      r8, r8
       mov      rcx, rdx
       mov      rdx, rax
       call     [System.SpanHelpers:Memmove(byref,byref,nuint)]
       nop
G_M28613_IG03:
       add      rsp, 40
       ret
G_M28613_IG04:
       call     [System.ThrowHelper:ThrowArgumentException_DestinationTooShort()]
       int3

; Total bytes of code 50

After:

G_M28613_IG01:
       sub      rsp, 40
G_M28613_IG02:
       mov      rax, bword ptr [rdx]
       mov      edx, dword ptr [rdx+0x08]
       mov      r10, bword ptr [rcx]
       mov      ecx, dword ptr [rcx+0x08]
       cmp      edx, ecx
       jg       SHORT G_M28613_IG06
       mov      r8d, edx
       add      r8, r8
       cmp      r8, 22
       je       SHORT G_M28613_IG04
G_M28613_IG03:
       mov      rcx, r10
       mov      rdx, rax
       call     [System.SpanHelpers:Memmove(byref,byref,nuint)]
       jmp      SHORT G_M28613_IG05
G_M28613_IG04:
       vmovdqu  xmm0, xmmword ptr [rax]
       vmovdqu  xmm1, xmmword ptr [rax+0x06]
       vmovdqu  xmmword ptr [r10], xmm0
       vmovdqu  xmmword ptr [r10+0x06], xmm1
G_M28613_IG05:
       add      rsp, 40
       ret
G_M28613_IG06:
       call     [System.ThrowHelper:ThrowArgumentException_DestinationTooShort()]
       int3

; Total bytes of code 78

Benchmark

Results from EgorBot/Benchmarks#600. Speedup is main / PR; values above 1.10X are bolded, and values below 1.00X indicate regressions. Length is in characters.

Length Apple M1 (macOS, ARM64) Cobalt 100 (Linux, ARM64) AMD EPYC 9V45 (Linux, x64)
2 2.58X 2.04X 2.51X
5 1.83X 1.74X 1.85X
16 1.86X 1.77X 1.04X
20 1.92X 1.59X 1.64X
32 1.75X 1.51X 2.10X
50 0.95X 1.00X 2.00X
64 0.97X 0.99X 1.89X
public class Bench {
    private string _src = null!;
    private char[] _dst = null!;

    [Params(2, 5, 16, 20, 32, 50, 64)]
    public int Length { get; set; }

    [GlobalSetup]
    public void Setup() {
        _src = new string('x', Length);
        _dst = new char[Length];
    }

    [Benchmark]
    public void Copy() => _src.AsSpan().CopyTo(_dst);
}

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 853b5c9d-0ee6-4e31-9290-82d2a20aa665
Copilot AI lite review requested due to automatic review settings September 17, 2026 22:20
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Add targeted regression coverage and preserve unrolling for constant-length calls.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Updates RyuJIT value profiling for Memmove and SequenceEqual in optimized instrumented tiers, including inline candidates.

Changes:

  • Adds value-profile metadata for both intrinsic paths.
  • Handles inline-candidate instrumentation.
  • Supports PGO-driven unrolling with default tiering and ReadyToRun.
File summaries
File Description
src/coreclr/jit/importercalls.cpp Adds value-profile handling for instrumented intrinsic calls.
Review details

Suppressed comments (1)

src/coreclr/jit/importercalls.cpp:1560

  • This now profiles every Memmove/SequenceEqual call in an optimized instrumented compilation, including calls whose length is already a constant. The value-probe inserter rewrites argument 2 to a comma containing CORINFO_HELP_VALUEPROFILE (fgprofile.cpp:2304-2327), while LowerCallMemmove/LowerCallMemcmp only unroll when that argument is still an integral constant (lower.cpp:2462-2468, 2552-2559). As a result, constant-size copies and comparisons lose their existing unrolling in the instrumented tier; avoid instrumenting constant lengths, and make the schema/visitor safe when a flagged block contains both constant and nonconstant calls.
    if (opts.IsInstrumented() && JitConfig.JitProfileValues() && call->IsCall() && call->AsCall()->IsSpecialIntrinsic())
    {
        const NamedIntrinsic ni = lookupNamedIntrinsic(call->AsCall()->gtCallMethHnd);
        if ((ni == NI_System_SpanHelpers_Memmove) || (ni == NI_System_SpanHelpers_SequenceEqual))
  • Files reviewed: 1/1 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread src/coreclr/jit/importercalls.cpp
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 853b5c9d-0ee6-4e31-9290-82d2a20aa665
Copilot AI review requested due to automatic review settings September 17, 2026 23:43
@EgorBo

This comment was marked as outdated.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Three unresolved moderate issues require fixes before approval.

Review details

Suppressed comments (3)

src/coreclr/jit/importercalls.cpp:5546

  • Can this new getHelperFtn query be made compatible with existing SuperPMI contexts? MethodContext::repGetHelperFtn requires a recorded entry and fails when the key is missing (src/coreclr/tools/superpmi/superpmi-shared/methodcontext.cpp:2355-2367), while contexts captured before this change will not contain a CORINFO_HELP_MEMCPY query for this path. Please use the established replay-compatibility mechanism or construct the existing helper without introducing an unconditional new EE query.
                CORINFO_METHOD_HANDLE memmoveHnd = NO_METHOD_HANDLE;
                info.compCompHnd->getHelperFtn(CORINFO_HELP_MEMCPY, nullptr, &memmoveHnd);
                if (memmoveHnd == NO_METHOD_HANDLE)

src/coreclr/jit/importercalls.cpp:1560

  • The existing Memmove/SequenceEqual tests cover semantics and constant-length unrolling, but the new behavior is the value-histogram schema/probe path in optimized instrumented tiers, including inline candidates. The PGO smoke test does not exercise either operation, so a regression in the new BBF_HAS_VALUE_PROFILE or inline-candidate handoff would pass; add a focused PGO test that warms variable lengths and verifies the profiled unroll or equivalent instrumentation result.
    // Collect value profiles in optimized instrumented tiers too, before wrapping inline candidates.
    if (opts.IsInstrumented() && JitConfig.JitProfileValues() && call->IsCall() && call->AsCall()->IsSpecialIntrinsic())
    {
        const NamedIntrinsic ni = lookupNamedIntrinsic(call->AsCall()->gtCallMethHnd);
        if ((ni == NI_System_SpanHelpers_Memmove) || (ni == NI_System_SpanHelpers_SequenceEqual))

src/coreclr/jit/importercalls.cpp:3696

  • Recognizing NI_System_Buffer_Memmove adds a new call shape, but fgbasic.cpp still raises CALLSITE_UNROLLABLE_MEMOP only for the SpanHelpers Memmove/SequenceEqual intrinsics (src/coreclr/jit/fgbasic.cpp:1140-1153). A constant Buffer.Memmove element count is lowered to the same byte Memmove call here, so inline candidates can miss the unrolling profitability hint and lose the constant unroll this intrinsic enables. Add the Buffer intrinsic to that observation or otherwise preserve the hint.
            case NI_System_Buffer_Memmove:
  • Files reviewed: 2/2 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

Clarify generic memmove expansion and helper naming.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 853b5c9d-0ee6-4e31-9290-82d2a20aa665
Copilot AI review requested due to automatic review settings September 18, 2026 00:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

A critical shared-generic instantiation issue and a missing PGO regression test remain.

Get a fresh assessment by requesting another Copilot review.

Review details

Suppressed comments (2)

src/coreclr/jit/importercalls.cpp:5546

  • This introduces a deferred TODO without the issue-linked, searchable form required for tracked work. Please either remove the TODO or attach it to a tracking issue using the repository's prefixed convention (for example, TODO-JIT:).
                // TODO: Rename CORINFO_HELP_MEMCPY to CORINFO_HELP_MEMMOVE to reflect its overlap-safe semantics.

src/coreclr/jit/importercalls.cpp:1560

  • This new path is only exercised when TieredPGO produces an instrumented optimized tier, including the case where the Memmove/SequenceEqual call is an inline candidate. The existing memmove/sequence-equal tests cover constant or semantic unrolling, while InstrumentedTiers is only a generic smoke test; none verifies that the value histogram is collected and then consumed for the optimized tier. Please add a PGO regression test that forces promotion and exercises both the variable-length fallback and the profile-driven unrolled path.
    if (opts.IsInstrumented() && JitConfig.JitProfileValues() && call->IsCall() && call->AsCall()->IsSpecialIntrinsic())
    {
        const NamedIntrinsic ni = lookupNamedIntrinsic(call->AsCall()->gtCallMethHnd);
        if ((ni == NI_System_SpanHelpers_Memmove) || (ni == NI_System_SpanHelpers_SequenceEqual))
  • Files reviewed: 3/3 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread src/coreclr/jit/importercalls.cpp
@EgorBo

EgorBo commented Sep 18, 2026

Copy link
Copy Markdown
Member Author

@EgorBot -macos_arm -linux_arm -linux_amd --envvars DOTNET_JitDisasm:Copy

using System;
using BenchmarkDotNet.Attributes;

public class Bench {
    private string _src = null!;
    private char[] _dst = null!;

    [Params(2, 5, 16, 20, 32, 50, 64)]
    public int Length { get; set; }

    [GlobalSetup]
    public void Setup() {
        _src = new string('x', Length);
        _dst = new char[Length];
    }

    [Benchmark]
    public void Copy() => _src.AsSpan().CopyTo(_dst);
}

Note

Benchmark snippet generated with GitHub Copilot.

@EgorBo

EgorBo commented Sep 18, 2026 •

Copy link
Copy Markdown
Member Author

PTAL @AndyAyersMS @dotnet/jit-contrib

This PGO-driven memmove finally works (after Andy's PR to instrument inlinees + this PR).

Normally, Span.CopyTo calls Buffer.Memmove<T>(T&,T&) and that one calls Buffer.Memmove(byte&,byte&). The latter is intrinsified, but because it's called from many places its profile is polluted, so I ended up intrinsifying Buffer.Memmove<T>(T&,T&) (for primitives) and lower as a call to memmove helper which is unroll friendly. We might remove this hack if we get context-dependent PGO.

PS: I'll file a PR to rename CORINFO_HELP_MEMCPY into CORINFO_HELP_MEMMOVE for clarity since all possible impl of it guarantee memmove semantics for overlapped pointers.

SPMI and MihaBot aren't showing diffs because this needs real PGO profile.

@EgorBo
EgorBo requested a review from AndyAyersMS September 18, 2026 14:25
@EgorBo
EgorBo merged commit 9f1dae9 into dotnet:main Sep 18, 2026
137 of 140 checks passed
@EgorBo
EgorBo deleted the jit-profile-values-r2r branch September 18, 2026 22:51
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 19, 2026
@EgorBo

EgorBo commented Sep 22, 2026

Copy link
Copy Markdown
Member Author

EgorBo added a commit that referenced this pull request Sep 23, 2026
Since #134160 led to **69
benchmarks** improved, I decided to do the same for memcmp idiom.

### Benchmark

```cs
using BenchmarkDotNet.Attributes;

public class Bench
{
    private byte[] _bytes1 = null!;
    private byte[] _bytes2 = null!;

    [Params(1, 2, 5, 8, 16, 20, 32, 50, 64, 128)]
    public int Length { get; set; }

    [GlobalSetup]
    public void Setup()
    {
        _bytes1 = new byte[Length];
        _bytes2 = new byte[Length];
    }

    [Benchmark]
    public bool Bytes() => _bytes1.AsSpan().SequenceEqual(_bytes2);
}
```

Results: EgorBot/Benchmarks#612 (arm64 only
handles up to 32 bytes today).

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e9c37cab-1590-4d11-a950-db4e7f065a3d
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants