Description
Compiler::optAssertionProp dispatches GT_HWINTRINSIC to optAssertionProp_HWIntrinsic
unconditionally, and that helper dereferences comp->vnStore:
https://github.com/dotnet/runtime/blob/main/src/coreclr/jit/assertionprop.cpp#L5964-L5968
#if defined(FEATURE_HW_INTRINSICS)
case GT_HWINTRINSIC:
optAssertionProp_HWIntrinsic(this, tree->AsHWIntrinsic());
return nullptr;
#endif // FEATURE_HW_INTRINSICS
ValueNum op1VN = comp->vnStore->VNConservativeNormalValue(op1->gtVNPair);
But optAssertionProp also runs during local assertion propagation in morph, which happens
before value numbering — vnStore is still null there. The method's own header comment notes this
(stmt may be nullptr during local assertion prop), and the neighbouring GT_ARR_LENGTH case
guards itself accordingly:
case GT_ARR_LENGTH:
...
if (!optLocalAssertionProp)
{
return optAssertionProp_Ind(assertions, tree, stmt);
}
return nullptr;
The GT_HWINTRINSIC case has no such guard, so local assertion prop calls
optAssertionProp_HWIntrinsic with vnStore == nullptr.
This is arm64-specific in practice because the helper only acts on
NI_Vector_ExtractMostSignificantBits, introduced for arm64 in #129688.
Regression
Introduced by #129688 (23ec3d982ce25e7b3037296eeed2dec0d0837795, 2026-08-04), which added
optAssertionProp_HWIntrinsic and the unguarded GT_HWINTRINSIC dispatch in the same change. The
helper is byte-identical on main today, so nothing has moved or modified it since.
It went unnoticed because the only consumer that runs an instrumented JIT is JIT PGO data
collection, and that pipeline was pinned to a pre-#129688 SDK until recently:
| When |
What |
| 2026-07-07 |
VMR ca60c43d982bb8e171b66ef8a38a777e004768ba, behind SDK 11.0.100-preview.7.26357.109 |
| 2026-08-04 |
#129688 merges |
| 2026-09-09 |
Last green arm64 PGO training run, still on the July 7 JIT |
| 2026-09-15 |
SDK pin moves to 11.0.100-rc.2.26461.115, the first pin containing #129688 |
The ExtractMostSignificantBits test added alongside #129688 does not catch this, because a
release JIT optimizes the null load away and the pass silently does nothing.
Why this is usually invisible
In an optimized (shipping) JIT build the null load is hidden: the compiler reorders around an
early-out so the null vnStore field is never actually loaded, and the pass silently does nothing.
The dereference is still UB, and any codegen change can expose it.
It does crash today in an instrumented (LLVM profile-generate) build of the JIT — the build
used to collect JIT PGO training data. There, the field load is emitted unconditionally and the
process dies with SIGSEGV.
Repro
Built and run on linux-arm64 with tiered compilation disabled, using
dotnet-sdk-pgo-11.0.100-rc.2.26470.103-linux-arm64 (the instrumented-JIT SDK published at
https://ci.dot.net/public/Sdk/11.0.100-rc.2.26470.103/):
[MethodImpl(MethodImplOptions.NoInlining)]
static uint Guarded128(object guard, Vector128<byte> input, byte threshold)
{
if (guard == null) throw new ArgumentNullException(nameof(guard));
Vector128<byte> mask = Vector128.GreaterThanOrEqual(input, Vector128.Create(threshold));
return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}
(The second use of mask is what keeps it a single-def GT_LCL_VAR operand of the
ExtractMostSignificantBits node, which is what the helper looks for.)
DOTNET_TieredCompilation=0 dotnet exec VectorMaskProbe.dll guarded128
The full program is below; it has four independent shapes
(guarded128, array128, guarded64, guarded-twice). All four crash. Results:
| JIT build |
DOTNET_EnableHWIntrinsic |
Result |
instrumented (dotnet-sdk-pgo) |
1 (default) |
SIGSEGV, all 4 cases |
instrumented (dotnet-sdk-pgo) |
0 |
all 4 pass |
ordinary (dotnet-sdk) |
1 (default) |
all 4 pass |
ordinary (dotnet-sdk) |
0 |
all 4 pass |
Same SDK version and same VMR commit (68f63657247e6ddb48a93cae3dfc05776dec8838) for both the
ordinary and instrumented SDK, on the same machine — so the only variable is how the JIT itself was
compiled.
I expect a checked JIT to fault (or assert) on the same program, since the null load is not
optimized away there, but I have not run that configuration myself.
Full program (VectorMaskProbe.cs)
using System;
using System.Numerics;
using System.Runtime.CompilerServices;
using System.Runtime.InteropServices;
using System.Runtime.Intrinsics;
Console.WriteLine(RuntimeInformation.FrameworkDescription);
Console.WriteLine($"Architecture: {RuntimeInformation.ProcessArchitecture}; Vector128 accelerated: {Vector128.IsHardwareAccelerated}");
Console.WriteLine($"CASE: {args[0]}");
Console.Out.Flush();
var data = new byte[32];
data[0] = 255;
var value = Vector128.LoadUnsafe(ref data[0]);
uint result;
uint expected;
switch (args[0])
{
case "guarded128":
result = Guarded128(new object(), value, 128);
expected = 256;
break;
case "array128":
result = Array128(data, 0, 128);
expected = 256;
break;
case "guarded64":
result = Guarded64(new object(), value.GetLower(), 128);
expected = 256;
break;
case "guarded-twice":
result = GuardedTwice(new object(), value, 128);
expected = 1;
break;
default:
throw new ArgumentException("Unknown probe case");
}
if (result != expected)
throw new InvalidOperationException($"Expected {expected}, got {result}");
Console.WriteLine($"PASS: {args[0]} = {result}");
[MethodImpl(MethodImplOptions.NoInlining)]
static uint Guarded128(object guard, Vector128<byte> input, byte threshold)
{
if (guard == null) throw new ArgumentNullException(nameof(guard));
Vector128<byte> mask = Vector128.GreaterThanOrEqual(input, Vector128.Create(threshold));
return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}
[MethodImpl(MethodImplOptions.NoInlining)]
static uint Array128(byte[] array, int offset, byte threshold)
{
if (array == null) throw new ArgumentNullException(nameof(array));
if ((uint)offset > (uint)(array.Length - 16)) throw new ArgumentOutOfRangeException(nameof(offset));
Vector128<byte> value = Vector128.LoadUnsafe(ref array[offset]);
Vector128<byte> mask = Vector128.GreaterThanOrEqual(value, Vector128.Create(threshold));
return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}
[MethodImpl(MethodImplOptions.NoInlining)]
static uint Guarded64(object guard, Vector64<byte> input, byte threshold)
{
if (guard == null) throw new ArgumentNullException(nameof(guard));
Vector64<byte> mask = Vector64.GreaterThanOrEqual(input, Vector64.Create(threshold));
return mask.ExtractMostSignificantBits() + mask.GetElement(0);
}
[MethodImpl(MethodImplOptions.NoInlining)]
static uint GuardedTwice(object guard, Vector128<byte> input, byte threshold)
{
if (guard == null) throw new ArgumentNullException(nameof(guard));
Vector128<byte> mask = Vector128.GreaterThanOrEqual(input, Vector128.Create(threshold));
return (uint)(BitOperations.PopCount(mask.ExtractMostSignificantBits()) +
BitOperations.TrailingZeroCount(mask.ExtractMostSignificantBits()));
}
It was compiled straight to IL with csc against the SDK's reference pack — no restore, no build
server — and each case was run as its own process, so nothing about the build differs between the
two runtimes.
Native evidence
GDB on the instrumented JIT, libclrjit.so load base 0xee2f513e0000:
Thread 1 "dotnet" received signal SIGSEGV
#0 0x0000ee2f514c1f90 libclrjit.so + 0xe1f90
#1 0x0000ee2f516d4080 libclrjit.so + 0x2f4080
#2 0x0000ee2f516dfc0c libclrjit.so + 0x2ffc0c
#3 0x0000ee2f516d4104 libclrjit.so + 0x2f4104
...
0xee2f514c1f84: ldr x9, [x26, #792] // x26 = Compiler*, +0x318 = vnStore
0xee2f514c1f88: ldr x1, [x8, #16]
0xee2f514c1f8c: mov w2, #0x1 // VNK_Conservative
=> 0xee2f514c1f90: ldr x0, [x9, #280] // x9 == 0
0xee2f514c1f94: bl 0xee2f517969fc
x1 0xffffffffffffffff -1 // NoVN
x9 0x0 0 // vnStore
The shipped binary is stripped, so I mapped those offsets back to functions using the binary's own
__llvm_prf_names / __llvm_prf_data records and .eh_frame ranges (this is an instrumented
build, so every function carries its own name and counter addresses):
| Offset |
Function |
+0xe1f90 |
Compiler::optAssertionProp (with optAssertionProp_HWIntrinsic inlined into it) |
+0x2f4080, +0x2f4104 |
Compiler::fgMorphTree |
+0x2ffc0c |
Compiler::fgMorphSmpOp |
+0x3b69fc (call target) |
ValueNumStore::VNNormalValue |
fgMorphTree → fgMorphSmpOp → optAssertionProp is the local assertion prop path in morph,
which confirms the pass is running before value numbering.
Suggested fix
Guard the case the same way GT_ARR_LENGTH is guarded:
#if defined(FEATURE_HW_INTRINSICS)
case GT_HWINTRINSIC:
// optAssertionProp_HWIntrinsic reads value numbers, which do not exist yet during local
// assertion prop: vnStore is not created until fgValueNumber runs.
if (!optLocalAssertionProp)
{
optAssertionProp_HWIntrinsic(this, tree->AsHWIntrinsic());
}
return nullptr;
#endif // FEATURE_HW_INTRINSICS
optAssertionProp_HWIntrinsic has exactly one caller, so this is the only site that needs it. A
vnStore != nullptr check would also work, but !optLocalAssertionProp matches the existing
convention in this switch and states the actual precondition.
This only disables an optimization that could never have run correctly during local assertion prop
anyway — global assertion prop, where vnStore exists, is unaffected.
I have not opened a PR — happy to, or to leave the choice of guard to the JIT team.
Impact
This is currently blocking dotnet-optimization arm64 runs. Every Linux arm64 PGO training leg
dies the first time the instrumented JIT compiles a method with this shape, so no arm64 JIT profile
data can be collected for .NET 11 and the shipped arm64 JIT cannot be rebuilt with fresh profiles.
Windows arm64, Linux arm64 IBC, and every x64 leg are unaffected and green — this is the only
remaining failure.
Configuration
11.0.100-rc.2.26470.103, VMR 68f63657247e6ddb48a93cae3dfc05776dec8838
- linux-arm64, Azure Linux 3.0, 64 KiB pages
- Present on
main as of d4b2951c1fa53ed95aa5bea8ad8e8a0e0a5a2fe7
Description
Compiler::optAssertionPropdispatchesGT_HWINTRINSICtooptAssertionProp_HWIntrinsicunconditionally, and that helper dereferences
comp->vnStore:https://github.com/dotnet/runtime/blob/main/src/coreclr/jit/assertionprop.cpp#L5964-L5968
ValueNum op1VN = comp->vnStore->VNConservativeNormalValue(op1->gtVNPair);But
optAssertionPropalso runs during local assertion propagation in morph, which happensbefore value numbering —
vnStoreis still null there. The method's own header comment notes this(
stmt may be nullptr during local assertion prop), and the neighbouringGT_ARR_LENGTHcaseguards itself accordingly:
The
GT_HWINTRINSICcase has no such guard, so local assertion prop callsoptAssertionProp_HWIntrinsicwithvnStore == nullptr.This is arm64-specific in practice because the helper only acts on
NI_Vector_ExtractMostSignificantBits, introduced for arm64 in #129688.Regression
Introduced by #129688 (
23ec3d982ce25e7b3037296eeed2dec0d0837795, 2026-08-04), which addedoptAssertionProp_HWIntrinsicand the unguardedGT_HWINTRINSICdispatch in the same change. Thehelper is byte-identical on
maintoday, so nothing has moved or modified it since.It went unnoticed because the only consumer that runs an instrumented JIT is JIT PGO data
collection, and that pipeline was pinned to a pre-#129688 SDK until recently:
ca60c43d982bb8e171b66ef8a38a777e004768ba, behind SDK11.0.100-preview.7.26357.10911.0.100-rc.2.26461.115, the first pin containing #129688The
ExtractMostSignificantBitstest added alongside #129688 does not catch this, because arelease JIT optimizes the null load away and the pass silently does nothing.
Why this is usually invisible
In an optimized (shipping) JIT build the null load is hidden: the compiler reorders around an
early-out so the null
vnStorefield is never actually loaded, and the pass silently does nothing.The dereference is still UB, and any codegen change can expose it.
It does crash today in an instrumented (LLVM profile-generate) build of the JIT — the build
used to collect JIT PGO training data. There, the field load is emitted unconditionally and the
process dies with SIGSEGV.
Repro
Built and run on linux-arm64 with tiered compilation disabled, using
dotnet-sdk-pgo-11.0.100-rc.2.26470.103-linux-arm64(the instrumented-JIT SDK published athttps://ci.dot.net/public/Sdk/11.0.100-rc.2.26470.103/):(The second use of
maskis what keeps it a single-defGT_LCL_VARoperand of theExtractMostSignificantBitsnode, which is what the helper looks for.)The full program is below; it has four independent shapes
(
guarded128,array128,guarded64,guarded-twice). All four crash. Results:DOTNET_EnableHWIntrinsicdotnet-sdk-pgo)dotnet-sdk-pgo)dotnet-sdk)dotnet-sdk)Same SDK version and same VMR commit (
68f63657247e6ddb48a93cae3dfc05776dec8838) for both theordinary and instrumented SDK, on the same machine — so the only variable is how the JIT itself was
compiled.
I expect a checked JIT to fault (or assert) on the same program, since the null load is not
optimized away there, but I have not run that configuration myself.
Full program (
VectorMaskProbe.cs)It was compiled straight to IL with
cscagainst the SDK's reference pack — no restore, no buildserver — and each case was run as its own process, so nothing about the build differs between the
two runtimes.
Native evidence
GDB on the instrumented JIT,
libclrjit.soload base0xee2f513e0000:The shipped binary is stripped, so I mapped those offsets back to functions using the binary's own
__llvm_prf_names/__llvm_prf_datarecords and.eh_frameranges (this is an instrumentedbuild, so every function carries its own name and counter addresses):
+0xe1f90Compiler::optAssertionProp(withoptAssertionProp_HWIntrinsicinlined into it)+0x2f4080,+0x2f4104Compiler::fgMorphTree+0x2ffc0cCompiler::fgMorphSmpOp+0x3b69fc(call target)ValueNumStore::VNNormalValuefgMorphTree→fgMorphSmpOp→optAssertionPropis the local assertion prop path in morph,which confirms the pass is running before value numbering.
Suggested fix
Guard the case the same way
GT_ARR_LENGTHis guarded:optAssertionProp_HWIntrinsichas exactly one caller, so this is the only site that needs it. AvnStore != nullptrcheck would also work, but!optLocalAssertionPropmatches the existingconvention in this switch and states the actual precondition.
This only disables an optimization that could never have run correctly during local assertion prop
anyway — global assertion prop, where
vnStoreexists, is unaffected.I have not opened a PR — happy to, or to leave the choice of guard to the JIT team.
Impact
This is currently blocking
dotnet-optimizationarm64 runs. Every Linux arm64 PGO training legdies the first time the instrumented JIT compiles a method with this shape, so no arm64 JIT profile
data can be collected for .NET 11 and the shipped arm64 JIT cannot be rebuilt with fresh profiles.
Windows arm64, Linux arm64 IBC, and every x64 leg are unaffected and green — this is the only
remaining failure.
Configuration
11.0.100-rc.2.26470.103, VMR68f63657247e6ddb48a93cae3dfc05776dec8838mainas ofd4b2951c1fa53ed95aa5bea8ad8e8a0e0a5a2fe7