Skip to content

JIT: null vnStore dereference in optAssertionProp_HWIntrinsic during local assertion prop (arm64) #134435

Description

@LoopedBard3

Description

Compiler::optAssertionProp dispatches GT_HWINTRINSIC to optAssertionProp_HWIntrinsic
unconditionally, and that helper dereferences comp->vnStore:

https://github.com/dotnet/runtime/blob/main/src/coreclr/jit/assertionprop.cpp#L5964-L5968

#if defined(FEATURE_HW_INTRINSICS)
        case GT_HWINTRINSIC:
            optAssertionProp_HWIntrinsic(this, tree->AsHWIntrinsic());
            return nullptr;
#endif // FEATURE_HW_INTRINSICS
    ValueNum op1VN = comp->vnStore->VNConservativeNormalValue(op1->gtVNPair);

But optAssertionProp also runs during local assertion propagation in morph, which happens
before value numbering — vnStore is still null there. The method's own header comment notes this
(stmt may be nullptr during local assertion prop), and the neighbouring GT_ARR_LENGTH case
guards itself accordingly:

        case GT_ARR_LENGTH:
            ...
            if (!optLocalAssertionProp)
            {
                return optAssertionProp_Ind(assertions, tree, stmt);
            }
            return nullptr;

The GT_HWINTRINSIC case has no such guard, so local assertion prop calls
optAssertionProp_HWIntrinsic with vnStore == nullptr.

This is arm64-specific in practice because the helper only acts on
NI_Vector_ExtractMostSignificantBits, introduced for arm64 in #129688.

Regression

Introduced by #129688 (23ec3d982ce25e7b3037296eeed2dec0d0837795, 2026-08-04), which added
optAssertionProp_HWIntrinsic and the unguarded GT_HWINTRINSIC dispatch in the same change. The
helper is byte-identical on main today, so nothing has moved or modified it since.

It went unnoticed because the only consumer that runs an instrumented JIT is JIT PGO data
collection, and that pipeline was pinned to a pre-#129688 SDK until recently:

When What
2026-07-07 VMR ca60c43d982bb8e171b66ef8a38a777e004768ba, behind SDK 11.0.100-preview.7.26357.109
2026-08-04 #129688 merges
2026-09-09 Last green arm64 PGO training run, still on the July 7 JIT
2026-09-15 SDK pin moves to 11.0.100-rc.2.26461.115, the first pin containing #129688

The ExtractMostSignificantBits test added alongside #129688 does not catch this, because a
release JIT optimizes the null load away and the pass silently does nothing.

Why this is usually invisible

In an optimized (shipping) JIT build the null load is hidden: the compiler reorders around an
early-out so the null vnStore field is never actually loaded, and the pass silently does nothing.
The dereference is still UB, and any codegen change can expose it.

It does crash today in an instrumented (LLVM profile-generate) build of the JIT — the build
used to collect JIT PGO training data. There, the field load is emitted unconditionally and the
process dies with SIGSEGV.

Repro

Built and run on linux-arm64 with tiered compilation disabled, using
dotnet-sdk-pgo-11.0.100-rc.2.26470.103-linux-arm64 (the instrumented-JIT SDK published at
https://ci.dot.net/public/Sdk/11.0.100-rc.2.26470.103/):

[MethodImpl(MethodImplOptions.NoInlining)]
static uint Guarded128(object guard, Vector128<byte> input, byte threshold)
{
    if (guard == null) throw new ArgumentNullException(nameof(guard));
    Vector128<byte> mask = Vector128.GreaterThanOrEqual(input, Vector128.Create(threshold));
    return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}

(The second use of mask is what keeps it a single-def GT_LCL_VAR operand of the
ExtractMostSignificantBits node, which is what the helper looks for.)

DOTNET_TieredCompilation=0 dotnet exec VectorMaskProbe.dll guarded128

The full program is below; it has four independent shapes
(guarded128, array128, guarded64, guarded-twice). All four crash. Results:

JIT build DOTNET_EnableHWIntrinsic Result
instrumented (dotnet-sdk-pgo) 1 (default) SIGSEGV, all 4 cases
instrumented (dotnet-sdk-pgo) 0 all 4 pass
ordinary (dotnet-sdk) 1 (default) all 4 pass
ordinary (dotnet-sdk) 0 all 4 pass

Same SDK version and same VMR commit (68f63657247e6ddb48a93cae3dfc05776dec8838) for both the
ordinary and instrumented SDK, on the same machine — so the only variable is how the JIT itself was
compiled.

I expect a checked JIT to fault (or assert) on the same program, since the null load is not
optimized away there, but I have not run that configuration myself.

Full program (VectorMaskProbe.cs)
using System;
using System.Numerics;
using System.Runtime.CompilerServices;
using System.Runtime.InteropServices;
using System.Runtime.Intrinsics;

Console.WriteLine(RuntimeInformation.FrameworkDescription);
Console.WriteLine($"Architecture: {RuntimeInformation.ProcessArchitecture}; Vector128 accelerated: {Vector128.IsHardwareAccelerated}");
Console.WriteLine($"CASE: {args[0]}");
Console.Out.Flush();

var data = new byte[32];
data[0] = 255;
var value = Vector128.LoadUnsafe(ref data[0]);
uint result;
uint expected;
switch (args[0])
{
    case "guarded128":
        result = Guarded128(new object(), value, 128);
        expected = 256;
        break;
    case "array128":
        result = Array128(data, 0, 128);
        expected = 256;
        break;
    case "guarded64":
        result = Guarded64(new object(), value.GetLower(), 128);
        expected = 256;
        break;
    case "guarded-twice":
        result = GuardedTwice(new object(), value, 128);
        expected = 1;
        break;
    default:
        throw new ArgumentException("Unknown probe case");
}

if (result != expected)
    throw new InvalidOperationException($"Expected {expected}, got {result}");
Console.WriteLine($"PASS: {args[0]} = {result}");

[MethodImpl(MethodImplOptions.NoInlining)]
static uint Guarded128(object guard, Vector128<byte> input, byte threshold)
{
    if (guard == null) throw new ArgumentNullException(nameof(guard));
    Vector128<byte> mask = Vector128.GreaterThanOrEqual(input, Vector128.Create(threshold));
    return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}

[MethodImpl(MethodImplOptions.NoInlining)]
static uint Array128(byte[] array, int offset, byte threshold)
{
    if (array == null) throw new ArgumentNullException(nameof(array));
    if ((uint)offset > (uint)(array.Length - 16)) throw new ArgumentOutOfRangeException(nameof(offset));
    Vector128<byte> value = Vector128.LoadUnsafe(ref array[offset]);
    Vector128<byte> mask = Vector128.GreaterThanOrEqual(value, Vector128.Create(threshold));
    return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}

[MethodImpl(MethodImplOptions.NoInlining)]
static uint Guarded64(object guard, Vector64<byte> input, byte threshold)
{
    if (guard == null) throw new ArgumentNullException(nameof(guard));
    Vector64<byte> mask = Vector64.GreaterThanOrEqual(input, Vector64.Create(threshold));
    return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}

[MethodImpl(MethodImplOptions.NoInlining)]
static uint GuardedTwice(object guard, Vector128<byte> input, byte threshold)
{
    if (guard == null) throw new ArgumentNullException(nameof(guard));
    Vector128<byte> mask = Vector128.GreaterThanOrEqual(input, Vector128.Create(threshold));
    return (uint)(BitOperations.PopCount(mask.ExtractMostSignificantBits()) +
                  BitOperations.TrailingZeroCount(mask.ExtractMostSignificantBits()));
}

It was compiled straight to IL with csc against the SDK's reference pack — no restore, no build
server — and each case was run as its own process, so nothing about the build differs between the
two runtimes.

Native evidence

GDB on the instrumented JIT, libclrjit.so load base 0xee2f513e0000:

Thread 1 "dotnet" received signal SIGSEGV
#0  0x0000ee2f514c1f90   libclrjit.so + 0xe1f90
#1  0x0000ee2f516d4080   libclrjit.so + 0x2f4080
#2  0x0000ee2f516dfc0c   libclrjit.so + 0x2ffc0c
#3  0x0000ee2f516d4104   libclrjit.so + 0x2f4104
...
   0xee2f514c1f84:  ldr  x9, [x26, #792]     // x26 = Compiler*, +0x318 = vnStore
   0xee2f514c1f88:  ldr  x1, [x8, #16]
   0xee2f514c1f8c:  mov  w2, #0x1            // VNK_Conservative
=> 0xee2f514c1f90:  ldr  x0, [x9, #280]      // x9 == 0
   0xee2f514c1f94:  bl   0xee2f517969fc

x1  0xffffffffffffffff  -1                   // NoVN
x9  0x0                 0                    // vnStore

The shipped binary is stripped, so I mapped those offsets back to functions using the binary's own
__llvm_prf_names / __llvm_prf_data records and .eh_frame ranges (this is an instrumented
build, so every function carries its own name and counter addresses):

Offset Function
+0xe1f90 Compiler::optAssertionProp (with optAssertionProp_HWIntrinsic inlined into it)
+0x2f4080, +0x2f4104 Compiler::fgMorphTree
+0x2ffc0c Compiler::fgMorphSmpOp
+0x3b69fc (call target) ValueNumStore::VNNormalValue

fgMorphTree → fgMorphSmpOp → optAssertionProp is the local assertion prop path in morph,
which confirms the pass is running before value numbering.

Suggested fix

Guard the case the same way GT_ARR_LENGTH is guarded:

#if defined(FEATURE_HW_INTRINSICS)
        case GT_HWINTRINSIC:
            // optAssertionProp_HWIntrinsic reads value numbers, which do not exist yet during local
            // assertion prop: vnStore is not created until fgValueNumber runs.
            if (!optLocalAssertionProp)
            {
                optAssertionProp_HWIntrinsic(this, tree->AsHWIntrinsic());
            }
            return nullptr;
#endif // FEATURE_HW_INTRINSICS

optAssertionProp_HWIntrinsic has exactly one caller, so this is the only site that needs it. A
vnStore != nullptr check would also work, but !optLocalAssertionProp matches the existing
convention in this switch and states the actual precondition.

This only disables an optimization that could never have run correctly during local assertion prop
anyway — global assertion prop, where vnStore exists, is unaffected.

I have not opened a PR — happy to, or to leave the choice of guard to the JIT team.

Impact

This is currently blocking dotnet-optimization arm64 runs. Every Linux arm64 PGO training leg
dies the first time the instrumented JIT compiles a method with this shape, so no arm64 JIT profile
data can be collected for .NET 11 and the shipped arm64 JIT cannot be rebuilt with fresh profiles.
Windows arm64, Linux arm64 IBC, and every x64 leg are unaffected and green — this is the only
remaining failure.

Configuration

  • 11.0.100-rc.2.26470.103, VMR 68f63657247e6ddb48a93cae3dfc05776dec8838
  • linux-arm64, Azure Linux 3.0, 64 KiB pages
  • Present on main as of d4b2951c1fa53ed95aa5bea8ad8e8a0e0a5a2fe7

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions