Skip to content

Preserve mask element granularity in SIMD transformations - #133708

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:tannergooding-simd-mask-granularity-fixes
Sep 15, 2026
Merged

tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:tannergooding-simd-mask-granularity-fixes

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Summary

Preserve mask producer/consumer element granularity during SIMD folding and lowering:

  • Prove blend masks are uniform at the consumer's element width before replacing sign-bit selection with AND/AND_NOT.
  • Recompute mask lane count when comparison fallback changes element type. Restrict direct constant-mask shortcuts to the zero/all-bits-set cases they implement.
  • Widen embedded broadcasts only when their scalar is constant. Preserve runtime broadcast load widths, including during PTESTM containment. Correct total mask lane accounting for wider vectors and duplicate constant bits with unsigned arithmetic.
  • Preserve SVE mask-to-vector conversions when producer and consumer element widths differ.

Add six focused xarch regression cases and a compact SVE regression for the shared folding guard. Existing valid same-width optimizations remain enabled.

Fixes #133548
Fixes #133552
Fixes #133555

Validation

  • Built Checked/Release x64 JITs, Checked cross-JITs, the merged regression runner, and the ARM64 SVE regression project.
  • Each of the six final xarch cases fails against the preserved original JIT and passes with the fixes on Checked and Release, with actual AVX512BW/VL execution. The final cases also pass with AVX disabled (SSE4.1-only). Broader AVX2, MinOpts, and portable software-fallback coverage passed before trimming redundant test variants.
  • SVE before/after cross-codegen confirms vector operations are retained for reinterpreted masks and the same-width predicate optimization remains. SVE runtime execution was unavailable on the x64 host.
  • Local BenchmarkDotNet checks retained identical instructions and code sizes for homogeneous blend (28 bytes), constant-broadcast select (44 bytes), and same-width comparison (33 bytes). Timings were noisy; no speedup claim. These measurements predate the rebase onto current main.
  • JIT formatting passed before publication.

Note

This PR description was drafted by GitHub Copilot.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 11, 2026 17:07
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 11, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

A runtime broadcast can still be widened unsafely through the blend rewrite path.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review tier: Lite
Findings: 1 High severity

New issues introduced by this change (1)
Severity Finding
High severity src/​coreclr/​jit/​gentree.cpp — Do not bypass broadcast validation for the rationalizer caller View comment
What changed in this PR

This PR preserves SIMD mask element granularity across xarch lowering and SVE folding.

Changes:

  • Tightens mask and comparison optimizations.
  • Preserves runtime broadcast widths.
  • Adds xarch and SVE regression coverage.
File Description
src/​tests/​JIT/​Regression_ro_2/​Runtime_133555.cs Tests reinterpreted comparison masks.
src/​tests/​JIT/​Regression_ro_2/​Runtime_133552.cs Tests runtime and constant broadcast behavior.
src/​tests/​JIT/​Regression_ro_2/​Runtime_133548.cs Tests blend mask granularity.
src/​tests/​JIT/​opt/​SVE/​PredicateInstructions.cs Tests SVE mask granularity preservation.
src/​coreclr/​jit/​lowerxarch.cpp Updates xarch mask and broadcast lowering.
src/​coreclr/​jit/​gentree.cpp Guards broadcast widening and SVE mask folding.

Comment thread src/coreclr/jit/gentree.cpp
Copilot AI review requested due to automatic review settings September 13, 2026 17:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

An unresolved moderate issue remains in vector-constant broadcast widening in gentree.cpp.

Review tier: Lite
Findings: None

Resolved findings (1)

@tannergooding

Copy link
Copy Markdown
Member Author

/ba-g unrelated and known issues

@tannergooding
tannergooding merged commit ae47da8 into dotnet:main Sep 15, 2026
140 of 143 checks passed
@tannergooding
tannergooding deleted the tannergooding-simd-mask-granularity-fixes branch September 15, 2026 01:42
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 16, 2026
jtschuster pushed a commit to jtschuster/runtime that referenced this pull request Sep 18, 2026
)

## Summary

Preserve mask producer/consumer element granularity during SIMD folding
and lowering:

- Prove blend masks are uniform at the consumer's element width before
replacing sign-bit selection with AND/AND_NOT.
- Recompute mask lane count when comparison fallback changes element
type. Restrict direct constant-mask shortcuts to the zero/all-bits-set
cases they implement.
- Widen embedded broadcasts only when their scalar is constant. Preserve
runtime broadcast load widths, including during PTESTM containment.
Correct total mask lane accounting for wider vectors and duplicate
constant bits with unsigned arithmetic.
- Preserve SVE mask-to-vector conversions when producer and consumer
element widths differ.

Add six focused xarch regression cases and a compact SVE regression for
the shared folding guard. Existing valid same-width optimizations remain
enabled.

Fixes dotnet#133548
Fixes dotnet#133552
Fixes dotnet#133555

## Validation

- Built Checked/Release x64 JITs, Checked cross-JITs, the merged
regression runner, and the ARM64 SVE regression project.
- Each of the six final xarch cases fails against the preserved original
JIT and passes with the fixes on Checked and Release, with actual
AVX512BW/VL execution. The final cases also pass with AVX disabled
(SSE4.1-only). Broader AVX2, MinOpts, and portable software-fallback
coverage passed before trimming redundant test variants.
- SVE before/after cross-codegen confirms vector operations are retained
for reinterpreted masks and the same-width predicate optimization
remains. SVE runtime execution was unavailable on the x64 host.
- Local BenchmarkDotNet checks retained identical instructions and code
sizes for homogeneous blend (28 bytes), constant-broadcast select (44
bytes), and same-width comparison (33 bytes). Timings were noisy; no
speedup claim. These measurements predate the rebase onto current main.
- JIT formatting passed before publication.

> [!NOTE]
> This PR description was drafted by GitHub Copilot.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@tannergooding

Copy link
Copy Markdown
Member Author

/backport to release/11.0

Note

This backport request was posted by GitHub Copilot at the user's direction.

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/11.0 (link to workflow run)

tannergooding added a commit that referenced this pull request Sep 20, 2026
…ons (#134213)

Backport of #133708 to release/11.0

/cc @tannergooding

## Customer Impact

- [ ] Customer reported
- [x] Found internally

This fixes incorrect SIMD results and JIT assertion failures in
#133548 and #133552. Both issues reproduce
on .NET 11 RC2 but not .NET 8, 9, or 10, so .NET 11 would otherwise ship
with regressions in SIMD mask folding and lowering. Because the fixes
share the same mask-granularity invariants, this backport also includes
#133555, which affects .NET 9 through .NET 11.

## Regression

- [x] Yes
- [ ] No

The #133548 and #133552 failures were
introduced during the .NET 11 development cycle; the exact introducing
change was not identified. #133555 predates .NET 11 and is
included because it is fixed by the same change.

## Testing

The source PR added focused regressions for all three xarch failures and
SVE coverage for the shared folding guard. The xarch cases failed with
the original JIT and passed after the fix in Checked and Release builds,
including native AVX-512 execution and reduced-ISA configurations. SVE
before/after cross-codegen preserved the required vector operations.
Representative valid same-width optimizations retained identical
instruction counts and code sizes.

## Risk

Low. The changes are narrowly scoped element-granularity guards and
invariant corrections that preserve existing valid SIMD transformations.
Focused regression coverage exercises the affected vector widths and ISA
configurations, SVE cross-codegen validates the shared guard, and
representative unaffected optimizations retained identical instruction
counts and code sizes.

**IMPORTANT**: If this backport is for a servicing release, please
verify that:

- For .NET 8 and .NET 9: The PR target branch is `release/X.0-staging`,
not `release/X.0`.
- For .NET 10+: The PR target branch is `release/X.0` (no `-staging`
suffix).

## Package authoring no longer needed in .NET 9

**IMPORTANT**: Starting with .NET 9, you no longer need to edit a NuGet
package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older
versions.

> [!NOTE]
> This backport template was completed by GitHub Copilot at the user's
direction.

Co-authored-by: Tanner Gooding <tagoo@outlook.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

3 participants