You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Improve both directions of struct field handling for incoming and outgoing ABI registers:
extract scalar fields directly from incoming parameter-register locals
reconstruct an 8-byte floating-point register directly from two promoted float fields
The urgent WASM correctness fix was split into #131237. This PR remains the broader JIT exploration for native code quality and the generalized implementation.
Changes
Use gtNewSimdGetElementNode for aligned scalar fields contained in SIMD parameter registers.
Reinterpret aligned native scalar floating-point carriers as SIMD values for direct lane extraction. The x64 double-to-second-float case emits movshdup; ARM64 emits dup.
Retain integer-carrier shift/narrow/bitcast handling for misaligned scalar fields and targets without the direct SIMD path.
Keep ARM32 on the existing lowering fallback because scalar extraction can require TYP_LONG nodes after long decomposition.
Synchronize Promotion::MapsToParameterRegister with lowering eligibility.
Teach xarch field-list lowering to reconstruct exactly two float fields in one 8-byte floating-point ABI register with NI_X86Base_UnpackLow, producing unpcklps instead of stack stores and reloads for arguments and returns.
Remove the experimental induced-promotion suppression heuristic; induced accesses are again costed like normal accesses.
Code generation
For a constructed { float X; float Y; } value:
Path
Before
After
Return
22 bytes / PerfScore 6.25
4 bytes / PerfScore 2.00
Argument forwarding
35 bytes / PerfScore 9.75
22 bytes / PerfScore 6.75
Both paths now use one unpcklps and avoid the spill/reload sequence.
Native Linux x64 BenchmarkDotNet results (AMD EPYC 7763):
Method
Stack baseline
insertps prototype
unpcklps
Return pair
8.286 ns
4.199 ns
2.737 ns
Construct + pass
9.402 ns
5.362 ns
4.384 ns
SuperPMI results
Latest CI run (build 1565991). Collections reporting code size differences:
This comes from synchronizing Promotion::MapsToParameterRegister with lowering. On ARM32, which has no FEATURE_SIMD, eligibility changes from float register with integer access at the segment offset to float register with floating-point access at the segment offset. Promotion therefore becomes more willing to split floating-point parameter registers into float replacements.
The float-pair reconstruction added here is TARGET_XARCH only, so ARM32 has no way to rebuild the aggregate and must spill and reload at P/Invoke argument sites. That is the same extract-then-reinsert shape @jakobbotsch described, on a target that lacks the backend support which makes it profitable on x64.
The options I see are to accept this, or to keep MapsToParameterRegister conservative on targets where the backend cannot reconstruct. The latter re-narrows the synchronization this PR is trying to achieve, so I would rather agree on the direction than pick one unilaterally.
Test infrastructure
JIT/Directed/StructABI/FieldListFloatInsertion was not actually running in CI, so its unpcklps assertions were never evaluated. Two separate causes:
JIT/Directed uses hand-maintained MergedWrapperProjectReference lists rather than the glob JIT/opt uses, and the new project was never added to one. No runner referenced it, so it appeared in neither the passed nor the skipped results of the Pri0 legs.
The project did not set RequiresProcessIsolation, so even once referenced it would have run in-process, and the generated run script that exports DOTNET_JitDisasm and invokes SuperFileCheck would never have been used.
Both are fixed, along with DOTNET_TieredCompilation=0 and DOTNET_JITMinOpts=0 to match the other disasm-check projects that use positive checks.
JIT/Directed/StructABI/FieldListFloatInsertion on osx-x64: SuperFileCheck resolves all three methods and the disasm output contains three unpcklps instructions
Same test on osx-arm64: the file declares no ARM64 patterns, so the method list is empty and the check skips cleanly rather than failing
Existing HFA suite with the Checked x64 JIT
Native Linux x64 matched-JIT BenchmarkDotNet comparison
The parameter-register-to-local mapping in lowering maps a scalar field of an
unenregisterable parameter to the local created for the containing register
segment, retyping the local access to the field's scalar type. On x64/arm64 a
scalar overlaps the low bytes of the vector register, so reading the register
local at scalar width is a free, correct reinterpret.
On wasm this is invalid: a v128 and a scalar (f64/f32/...) are distinct
value-stack types with no implicit reinterpret, so 'local.get <v128>' typed as
a scalar produces invalid wasm (wasm-tools: 'expected f64, found v128'). This
surfaces when crossgen2 compiles, e.g., Vector128<T>.GetHashCode (whose
GetElementUnsafe reads a scalar lane of the vector) inlined into
GenericEqualityComparer<Vector128<double>>.GetHashCode.
Guard the mapping on wasm: skip mapping a scalar field to a SIMD register
segment so the access falls back to the (memory-homed) unenregisterable
parameter and a normal scalar load is emitted. This is the floating-point
ToScalar / SIMD->scalar var_type case noted as deferred in dotnet#130444.
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
This PR adjusts JIT lowering for the wasm target to avoid inducing an invalid scalar read from a SIMD (v128) register-local when the accessed parameter field is a non-SIMD scalar type. The change is localized to the parameter-register-local mapping logic and is guarded to affect only wasm.
Changes:
Add a TARGET_WASM-guarded check in Lowering::FindInducedParameterRegisterLocals to skip mapping when the ABI segment is SIMD but the field access is scalar.
Preserve existing behavior for non-wasm targets and for SIMD-typed field accesses.
Yes — this is effectively on top of #130866 (it's in a wasm R2R prototype branch that has #130866 merged, which materializes `Vector128` as a wasm `v128` parameter — hence the v128 register parameter here). Agreed on `f64x2.extract_lane 0` (and the general extract/insert for the sibling paths) being the better, spill-free lowering. Let me look at doing it that way in this PR rather than the memory fallback.
Note
This comment was authored with the assistance of GitHub Copilot.
Per review feedback (thanks @tannergooding, @AndyAyersMS): rather than skipping the
parameter-register-local mapping for a scalar field of a SIMD register (which forces a spill/reload
through the memory-homed parameter), extract the lane directly from the SIMD register local via
gtNewSimdGetElementNode. This lowers to an explicit lane extract (e.g. f64x2.extract_lane 0) with no
spill. The GetElement node is lowered by the subsequent per-block lowering pass.
> [!NOTE]
> This change was authored with the assistance of GitHub Copilot.
Use direct SIMD lane extraction and scalar bit manipulation for fields carried in floating-point parameter registers. Keep physical promotion eligibility synchronized and retain the ARM32 fallback.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 22214765-2cd2-4951-b778-8793730f2bc1
lewing
changed the title
JIT: Fix wasm scalar-field extraction from a SIMD-register parameter
JIT: Improve parameter register field extraction
Jul 22, 2026
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.
Suppressed comments (2)
src/tests/JIT/opt/Unsafe/Unsafe.cs:75
This test constructs a double containing two float bit patterns and then reads the “second float” via Unsafe.As<double, float> + Unsafe.Add. That interpretation depends on the platform endianness, so it can produce different results (and false failures) on non-little-endian targets. Consider skipping these checks when BitConverter.IsLittleEndian is false (or computing the expected value based on endianness).
The project file enables disasm checking but doesn’t follow the established pattern used by other <HasDisasmCheck>true</HasDisasmCheck> tests (e.g., src/tests/JIT/opt/Unsafe/Unsafe.csproj) to force process isolation and disable tiered compilation/minopts. Without these, the disasm checks can become flaky or validate Tier0 codegen instead of the intended optimized codegen.
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 2 comments.
Suppressed comments (1)
src/tests/JIT/opt/Unsafe/Unsafe.cs:80
This test constructs and inspects bit patterns via Int64BitsToDouble/DoubleToInt64Bits shifting, which assumes little-endian layout for {float,float} within a double and for the byte-offset-2 misaligned read. On big-endian targets these expectations won’t match the actual memory layout. Consider building the double from a byte[]/Span composed from BitConverter.GetBytes(float) and computing expectedMisaligned via BitConverter.ToSingle(bytes, 2) to make the test endian-agnostic.
FieldListFloatInsertion did not set RequiresProcessIsolation, so it was
compiled into the merged test runner and executed in-process. The
generated run script that exports DOTNET_JitDisasm and invokes
SuperFileCheck was never used, leaving the unpcklps assertions dead.
Add RequiresProcessIsolation along with DOTNET_TieredCompilation=0 and
DOTNET_JITMinOpts=0, matching every other disasm-check project that uses
positive checks.
Also require FEATURE_HW_INTRINSICS alongside TARGET_XARCH for the
float-pair reconstruction blocks, since InsertNewSimdCreateScalarUnsafeNode
and gtNewSimdHWIntrinsicNode are only declared under that feature.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
JIT/Directed uses hand-maintained MergedWrapperProjectReference lists
rather than the glob that JIT/opt uses, and the new project was never
added to one. No runner referenced it, so the test never ran in CI in
any form - it appears in neither the passed nor the skipped results for
the Pri0 legs. Building it directly with -Test or -Dir hid this, since
those bypass runner discovery.
Add it to Directed_ro.csproj alongside the other StructABI projects;
DebugType None with Optimize true matches that runner. The generated
runner now registers it as an out-of-process test, the same shape as
EmptyStructs.cmd.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Accept both the legacy two-operand and VEX three-operand forms of unpcklps in the FieldListFloatInsertion checks.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
ARM32 promotion needs a conservative exclusion or a corresponding reconstruction path.
Review tier: Lite Findings: None
Previously missed findings (1)
In code that hasn't changed since last review
Avoid re-enabling costly float promotion on ARM32
src/coreclr/jit/promotion.cpp:3077
On ARM32 this condition now allows MapsToParameterRegister to promote a float field from a floating-point parameter-register segment at offset 0. The induced extraction is supported, but ARM32 has no matching field-list reconstruction path, so calls that forward or P/Invoke these promoted values spill and reload; the PR reports a net +314-byte ARM32 regression, including +270 bytes in RayTracer:GetNaturalColor. Please retain the previous conservative exclusion for float accesses on ARM32 (while still allowing the existing integer-at-offset-zero case), or provide an ARM32 reconstruction before enabling this promotion.
// The scalar extraction in lowering can require TYP_LONG nodes, which are not legal after decomposition.
if (genIsValidFloatReg(seg.GetRegister()) && (!varTypeUsesFloatReg(accessType) || (offset != seg.Offset)))
{
continue;
}
#endif // TARGET_ARM
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Improve both directions of struct field handling for incoming and outgoing ABI registers:
floatfieldsThe urgent WASM correctness fix was split into #131237. This PR remains the broader JIT exploration for native code quality and the generalized implementation.
Changes
gtNewSimdGetElementNodefor aligned scalar fields contained in SIMD parameter registers.double-to-second-floatcase emitsmovshdup; ARM64 emitsdup.TYP_LONGnodes after long decomposition.Promotion::MapsToParameterRegisterwith lowering eligibility.floatfields in one 8-byte floating-point ABI register withNI_X86Base_UnpackLow, producingunpcklpsinstead of stack stores and reloads for arguments and returns.Code generation
For a constructed
{ float X; float Y; }value:Both paths now use one
unpcklpsand avoid the spill/reload sequence.Native Linux x64 BenchmarkDotNet results (AMD EPYC 7763):
insertpsprototypeunpcklpsSuperPMI results
Latest CI run (build 1565991). Collections reporting code size differences:
libraries.crossgen2.linux.x64.checkedlibraries.crossgen2.osx.arm64.checkedsmoke_tests.nativeaot.linux.arm64.checkedsmoke_tests.nativeaot.windows.arm64.checkedlibraries_tests_no_tiered_compilation.run.linux.arm64.Releasecoreclr_tests.run.linux.arm.checkedlibraries_tests_no_tiered_compilation.run.linux.arm.ReleaseThe archived earlier analysis, including the
insertps-versus-unpcklpscomparison and representative diffs, is at https://gist.github.com/lewing/59292126500746b46abd297f609cfad4.Known ARM32 regression
ARM32 is net +314 bytes and is the one area I would like reviewer input on before merge.
RayTracer:GetNaturalColor(SceneObject,Vector,Vector,Vector,Scene)SIMDTests.Vector3InteropTests.PInvokeTest:nativeCall_PInvoke_Vector3ArgHFATest.TestMan:Average19_HFA02Runtime_128373:ProblematicBodyCamera:CreateThis comes from synchronizing
Promotion::MapsToParameterRegisterwith lowering. On ARM32, which has noFEATURE_SIMD, eligibility changes from float register with integer access at the segment offset to float register with floating-point access at the segment offset. Promotion therefore becomes more willing to split floating-point parameter registers intofloatreplacements.The float-pair reconstruction added here is
TARGET_XARCHonly, so ARM32 has no way to rebuild the aggregate and must spill and reload at P/Invoke argument sites. That is the same extract-then-reinsert shape @jakobbotsch described, on a target that lacks the backend support which makes it profitable on x64.The options I see are to accept this, or to keep
MapsToParameterRegisterconservative on targets where the backend cannot reconstruct. The latter re-narrows the synchronization this PR is trying to achieve, so I would rather agree on the direction than pick one unilaterally.Test infrastructure
JIT/Directed/StructABI/FieldListFloatInsertionwas not actually running in CI, so itsunpcklpsassertions were never evaluated. Two separate causes:JIT/Directeduses hand-maintainedMergedWrapperProjectReferencelists rather than the globJIT/optuses, and the new project was never added to one. No runner referenced it, so it appeared in neither the passed nor the skipped results of the Pri0 legs.RequiresProcessIsolation, so even once referenced it would have run in-process, and the generated run script that exportsDOTNET_JitDisasmand invokes SuperFileCheck would never have been used.Both are fixed, along with
DOTNET_TieredCompilation=0andDOTNET_JITMinOpts=0to match the other disasm-check projects that use positive checks.Validation
JIT/opt/Unsafe/Unsafeextraction regression tests, confirmed emittingmovshdupJIT/Directed/StructABI/FieldListFloatInsertionon osx-x64: SuperFileCheck resolves all three methods and the disasm output contains threeunpcklpsinstructionsDOTNET_EnableSSE41=0,DOTNET_EnableHWIntrinsic=0)unpcklpsloweringNote
This PR was authored with the assistance of GitHub Copilot.