Skip to content

JIT: Improve parameter register field extraction and reconstruction - #131174

Open
lewing wants to merge 19 commits into
dotnet:mainfrom
lewing:wasm-r2r-simd-scalar-extract
Open

lewing wants to merge 19 commits into
dotnet:mainfrom
lewing:wasm-r2r-simd-scalar-extract

Conversation

@lewing

@lewing lewing commented Jul 21, 2026 •

Copy link
Copy Markdown
Member

Summary

Improve both directions of struct field handling for incoming and outgoing ABI registers:

  • extract scalar fields directly from incoming parameter-register locals
  • reconstruct an 8-byte floating-point register directly from two promoted float fields

The urgent WASM correctness fix was split into #131237. This PR remains the broader JIT exploration for native code quality and the generalized implementation.

Changes

  • Use gtNewSimdGetElementNode for aligned scalar fields contained in SIMD parameter registers.
  • Reinterpret aligned native scalar floating-point carriers as SIMD values for direct lane extraction. The x64 double-to-second-float case emits movshdup; ARM64 emits dup.
  • Retain integer-carrier shift/narrow/bitcast handling for misaligned scalar fields and targets without the direct SIMD path.
  • Keep ARM32 on the existing lowering fallback because scalar extraction can require TYP_LONG nodes after long decomposition.
  • Synchronize Promotion::MapsToParameterRegister with lowering eligibility.
  • Teach xarch field-list lowering to reconstruct exactly two float fields in one 8-byte floating-point ABI register with NI_X86Base_UnpackLow, producing unpcklps instead of stack stores and reloads for arguments and returns.
  • Remove the experimental induced-promotion suppression heuristic; induced accesses are again costed like normal accesses.

Code generation

For a constructed { float X; float Y; } value:

Path Before After
Return 22 bytes / PerfScore 6.25 4 bytes / PerfScore 2.00
Argument forwarding 35 bytes / PerfScore 9.75 22 bytes / PerfScore 6.75

Both paths now use one unpcklps and avoid the spill/reload sequence.

Native Linux x64 BenchmarkDotNet results (AMD EPYC 7763):

Method Stack baseline insertps prototype unpcklps
Return pair 8.286 ns 4.199 ns 2.737 ns
Construct + pass 9.402 ns 5.362 ns 4.384 ns

SuperPMI results

Latest CI run (build 1565991). Collections reporting code size differences:

Collection Delta
libraries.crossgen2.linux.x64.checked −733
libraries.crossgen2.osx.arm64.checked −112
smoke_tests.nativeaot.linux.arm64.checked −28
smoke_tests.nativeaot.windows.arm64.checked −28
libraries_tests_no_tiered_compilation.run.linux.arm64.Release +8
coreclr_tests.run.linux.arm.checked +346
libraries_tests_no_tiered_compilation.run.linux.arm.Release −32

The archived earlier analysis, including the insertps-versus-unpcklps comparison and representative diffs, is at https://gist.github.com/lewing/59292126500746b46abd297f609cfad4.

Known ARM32 regression

ARM32 is net +314 bytes and is the one area I would like reviewer input on before merge.

Method Delta
RayTracer:GetNaturalColor(SceneObject,Vector,Vector,Vector,Scene) +270 (4.90%)
SIMDTests.Vector3InteropTests.PInvokeTest:nativeCall_PInvoke_Vector3Arg +48 (10.34%)
same method, second context +48 (10.30%)
HFATest.TestMan:Average19_HFA02 +4
Runtime_128373:ProblematicBody −22
Camera:Create −2

This comes from synchronizing Promotion::MapsToParameterRegister with lowering. On ARM32, which has no FEATURE_SIMD, eligibility changes from float register with integer access at the segment offset to float register with floating-point access at the segment offset. Promotion therefore becomes more willing to split floating-point parameter registers into float replacements.

The float-pair reconstruction added here is TARGET_XARCH only, so ARM32 has no way to rebuild the aggregate and must spill and reload at P/Invoke argument sites. That is the same extract-then-reinsert shape @jakobbotsch described, on a target that lacks the backend support which makes it profitable on x64.

The options I see are to accept this, or to keep MapsToParameterRegister conservative on targets where the backend cannot reconstruct. The latter re-narrows the synchronization this PR is trying to achieve, so I would rather agree on the direction than pick one unilaterally.

Test infrastructure

JIT/Directed/StructABI/FieldListFloatInsertion was not actually running in CI, so its unpcklps assertions were never evaluated. Two separate causes:

  • JIT/Directed uses hand-maintained MergedWrapperProjectReference lists rather than the glob JIT/opt uses, and the new project was never added to one. No runner referenced it, so it appeared in neither the passed nor the skipped results of the Pri0 legs.
  • The project did not set RequiresProcessIsolation, so even once referenced it would have run in-process, and the generated run script that exports DOTNET_JitDisasm and invokes SuperFileCheck would never have been used.

Both are fixed, along with DOTNET_TieredCompilation=0 and DOTNET_JITMinOpts=0 to match the other disasm-check projects that use positive checks.

Validation

  • Debug ARM64 JIT build and functional test
  • Debug and Checked x64 JIT builds
  • JIT/opt/Unsafe/Unsafe extraction regression tests, confirmed emitting movshdup
  • JIT/Directed/StructABI/FieldListFloatInsertion on osx-x64: SuperFileCheck resolves all three methods and the disasm output contains three unpcklps instructions
  • Same test on osx-arm64: the file declares no ARM64 patterns, so the method list is empty and the check skips cleanly rather than failing
  • Existing HFA suite with the Checked x64 JIT
  • Native Linux x64 matched-JIT BenchmarkDotNet comparison
  • ISA-toggle checks (DOTNET_EnableSSE41=0, DOTNET_EnableHWIntrinsic=0)
  • JIT formatting
  • Independent multi-model reviews of extraction, exact field-list shape validation, and unpcklps lowering

Note

This PR was authored with the assistance of GitHub Copilot.

The parameter-register-to-local mapping in lowering maps a scalar field of an
unenregisterable parameter to the local created for the containing register
segment, retyping the local access to the field's scalar type. On x64/arm64 a
scalar overlaps the low bytes of the vector register, so reading the register
local at scalar width is a free, correct reinterpret.

On wasm this is invalid: a v128 and a scalar (f64/f32/...) are distinct
value-stack types with no implicit reinterpret, so 'local.get <v128>' typed as
a scalar produces invalid wasm (wasm-tools: 'expected f64, found v128'). This
surfaces when crossgen2 compiles, e.g., Vector128<T>.GetHashCode (whose
GetElementUnsafe reads a scalar lane of the vector) inlined into
GenericEqualityComparer<Vector128<double>>.GetHashCode.

Guard the mapping on wasm: skip mapping a scalar field to a SIMD register
segment so the access falls back to the (memory-homed) unenregisterable
parameter and a normal scalar load is emitted. This is the floating-point
ToScalar / SIMD->scalar var_type case noted as deferred in dotnet#130444.
Copilot AI lite review requested due to automatic review settings July 21, 2026 22:31
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adjusts JIT lowering for the wasm target to avoid inducing an invalid scalar read from a SIMD (v128) register-local when the accessed parameter field is a non-SIMD scalar type. The change is localized to the parameter-register-local mapping logic and is guarded to affect only wasm.

Changes:

  • Add a TARGET_WASM-guarded check in Lowering::FindInducedParameterRegisterLocals to skip mapping when the ABI segment is SIMD but the field access is scalar.
  • Preserve existing behavior for non-wasm targets and for SIMD-typed field accesses.

Comment thread src/coreclr/jit/lower.cpp Outdated
@AndyAyersMS

Copy link
Copy Markdown
Member

Is this on top of #130866? Else how do we have a V128 parameter?

I agree with Tanner, would be nice to use f64x2.extract_lane 0 here or similar to avoid having to spill and reload.

@lewing

lewing commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

Yes — this is effectively on top of #130866 (it's in a wasm R2R prototype branch that has #130866 merged, which materializes `Vector128` as a wasm `v128` parameter — hence the v128 register parameter here). Agreed on `f64x2.extract_lane 0` (and the general extract/insert for the sibling paths) being the better, spill-free lowering. Let me look at doing it that way in this PR rather than the memory fallback.

Note

This comment was authored with the assistance of GitHub Copilot.

Per review feedback (thanks @tannergooding, @AndyAyersMS): rather than skipping the
parameter-register-local mapping for a scalar field of a SIMD register (which forces a spill/reload
through the memory-homed parameter), extract the lane directly from the SIMD register local via
gtNewSimdGetElementNode. This lowers to an explicit lane extract (e.g. f64x2.extract_lane 0) with no
spill. The GetElement node is lowered by the subsequent per-block lowering pass.

> [!NOTE]
> This change was authored with the assistance of GitHub Copilot.
Copilot AI review requested due to automatic review settings July 21, 2026 23:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 2 comments.

Comment thread src/coreclr/jit/lower.cpp Outdated
Comment thread src/coreclr/jit/lower.cpp Outdated
lewing added a commit to lewing/runtime that referenced this pull request Jul 22, 2026
…tnet#131174)

Extract the scalar lane directly from the SIMD register local instead of a memory spill/reload.
lewing added a commit to lewing/runtime that referenced this pull request Jul 22, 2026
…tnet#131174)

Extract the scalar lane directly from the SIMD register local instead of a memory spill/reload.
lewing added 2 commits July 21, 2026 20:10
Use direct SIMD lane extraction and scalar bit manipulation for fields carried in floating-point parameter registers. Keep physical promotion eligibility synchronized and retain the ARM32 fallback.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 22214765-2cd2-4951-b778-8793730f2bc1
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 22214765-2cd2-4951-b778-8793730f2bc1
Copilot AI review requested due to automatic review settings July 22, 2026 01:24
@lewing lewing changed the title JIT: Fix wasm scalar-field extraction from a SIMD-register parameter JIT: Improve parameter register field extraction Jul 22, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 3 out of 3 changed files in this pull request and generated no new comments.

Copilot AI review requested due to automatic review settings July 22, 2026 01:33
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

Copilot AI review requested due to automatic review settings August 18, 2026 23:07
@lewing lewing added this to the 12.0.0 milestone Aug 18, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.

Suppressed comments (2)

src/tests/JIT/opt/Unsafe/Unsafe.cs:75

  • This test constructs a double containing two float bit patterns and then reads the “second float” via Unsafe.As<double, float> + Unsafe.Add. That interpretation depends on the platform endianness, so it can produce different results (and false failures) on non-little-endian targets. Consider skipping these checks when BitConverter.IsLittleEndian is false (or computing the expected value based on endianness).
            double doubleValue = BitConverter.Int64BitsToDouble(
                ((long)BitConverter.SingleToInt32Bits(2.5f) << 32) | (uint)BitConverter.SingleToInt32Bits(1.25f));

src/tests/JIT/Directed/StructABI/FieldListFloatInsertion.csproj:11

  • The project file enables disasm checking but doesn’t follow the established pattern used by other <HasDisasmCheck>true</HasDisasmCheck> tests (e.g., src/tests/JIT/opt/Unsafe/Unsafe.csproj) to force process isolation and disable tiered compilation/minopts. Without these, the disasm checks can become flaky or validate Tier0 codegen instead of the intended optimized codegen.
<Project Sdk="Microsoft.NET.Sdk">
  <PropertyGroup>
    <DebugType>None</DebugType>
    <Optimize>True</Optimize>
  </PropertyGroup>
  <ItemGroup>
    <Compile Include="$(MSBuildProjectName).cs">
      <HasDisasmCheck>true</HasDisasmCheck>
    </Compile>
  </ItemGroup>
</Project>

Copilot AI review requested due to automatic review settings August 22, 2026 04:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated 2 comments.

Suppressed comments (1)

src/tests/JIT/opt/Unsafe/Unsafe.cs:80

  • This test constructs and inspects bit patterns via Int64BitsToDouble/DoubleToInt64Bits shifting, which assumes little-endian layout for {float,float} within a double and for the byte-offset-2 misaligned read. On big-endian targets these expectations won’t match the actual memory layout. Consider building the double from a byte[]/Span composed from BitConverter.GetBytes(float) and computing expectedMisaligned via BitConverter.ToSingle(bytes, 2) to make the test endian-agnostic.
            double doubleValue = BitConverter.Int64BitsToDouble(
                ((long)BitConverter.SingleToInt32Bits(2.5f) << 32) | (uint)BitConverter.SingleToInt32Bits(1.25f));
            if (UnsafeAsSecondFloat_Double(doubleValue) != 2.5f)
                return 0;

            float expectedMisaligned = BitConverter.Int32BitsToSingle((int)(BitConverter.DoubleToInt64Bits(doubleValue) >> 16));
            if (UnsafeAsMisalignedFloat_Double(doubleValue) != expectedMisaligned)

Comment thread src/coreclr/jit/lower.cpp Outdated
Comment thread src/tests/JIT/Directed/StructABI/FieldListFloatInsertion.csproj
FieldListFloatInsertion did not set RequiresProcessIsolation, so it was
compiled into the merged test runner and executed in-process. The
generated run script that exports DOTNET_JitDisasm and invokes
SuperFileCheck was never used, leaving the unpcklps assertions dead.
Add RequiresProcessIsolation along with DOTNET_TieredCompilation=0 and
DOTNET_JITMinOpts=0, matching every other disasm-check project that uses
positive checks.

Also require FEATURE_HW_INTRINSICS alongside TARGET_XARCH for the
float-pair reconstruction blocks, since InsertNewSimdCreateScalarUnsafeNode
and gtNewSimdHWIntrinsicNode are only declared under that feature.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings August 24, 2026 22:57

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.

JIT/Directed uses hand-maintained MergedWrapperProjectReference lists
rather than the glob that JIT/opt uses, and the new project was never
added to one. No runner referenced it, so the test never ran in CI in
any form - it appears in neither the passed nor the skipped results for
the Pri0 legs. Building it directly with -Test or -Dir hid this, since
those bypass runner discovery.

Add it to Directed_ro.csproj alongside the other StructABI projects;
DebugType None with Optimize true matches that runner. The generated
runner now registers it as an out-of-process test, the same shape as
EmptyStructs.cmd.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings August 25, 2026 00:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated no new comments.

Accept both the legacy two-operand and VEX three-operand forms of unpcklps in the FieldListFloatInsertion checks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 12, 2026 19:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

ARM32 promotion needs a conservative exclusion or a corresponding reconstruction path.

Review tier: Lite
Findings: None

Previously missed findings (1)

In code that hasn't changed since last review

Medium severity Avoid re-enabling costly float promotion on ARM32

src/​coreclr/​jit/​promotion.cpp:3077

On ARM32 this condition now allows MapsToParameterRegister to promote a float field from a floating-point parameter-register segment at offset 0. The induced extraction is supported, but ARM32 has no matching field-list reconstruction path, so calls that forward or P/Invoke these promoted values spill and reload; the PR reports a net +314-byte ARM32 regression, including +270 bytes in RayTracer:GetNaturalColor. Please retain the previous conservative exclusion for float accesses on ARM32 (while still allowing the existing integer-at-offset-zero case), or provide an ARM32 reconstruction before enabling this promotion.

Copilot AI review requested due to automatic review settings September 17, 2026 23:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical SIMD extraction safety and ARM32 regression issues remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 1 Medium severity

Open (2)

Comment thread src/coreclr/jit/lower.cpp
Comment on lines +9305 to +9307
if (varTypeIsSIMD(segment.GetRegisterType()) &&
(varTypeIsSIMD(fld) ? (fld->GetLclOffs() != segment.Offset)
: (((fld->GetLclOffs() - segment.Offset) % genTypeSize(fld)) != 0)))
Comment on lines +3117 to +3123
#ifdef TARGET_ARM
// The scalar extraction in lowering can require TYP_LONG nodes, which are not legal after decomposition.
if (genIsValidFloatReg(seg.GetRegister()) && (!varTypeUsesFloatReg(accessType) || (offset != seg.Offset)))
{
continue;
}
#endif // TARGET_ARM

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants