Skip to content

Fix VNNI blend merge operand preservation - #134364

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-vnni-blend-merge-operand
Sep 22, 2026
Merged

tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-vnni-blend-merge-operand

Conversation

@tannergooding

Copy link
Copy Markdown
Member

VNNI multiply-and-accumulate instructions read their destination as the accumulator. When merge-masked, inactive lanes also retain that destination. Folding BlendVariable(fallback, MultiplyWideningAndAdd(addend, left, right), mask) into a masked VNNI instruction therefore cannot preserve an independent fallback: initializing the accumulator overwrites it.

Mark the regular and saturating VNNI intrinsic entries as HW_Flag_RmwIntrinsic, including the newer integer variants. This lets the existing embedded-masking compatibility check reject the invalid containment and emit the VNNI operation followed by the blend. No special-case lowering or codegen is needed.

Add 12 regression cases covering 128/256/512-bit vectors, byte/short products, and saturating/non-saturating forms. Distinct fallback/addend values, zero products, and alternating comparison lanes isolate merge preservation. The repro helpers force optimized compilation even under default tiering.

Validation on Windows x64 with AVX-512 VNNI hardware:

  • Built Checked and Release JITs and the regression runner; ran JIT formatting.
  • All 12 cases fail with the original Checked JIT and pass with the fixed Checked JIT under default tiering. All 12 also pass with fixed Checked and Release JITs with tiering disabled.
  • Disassembly confirms actual VNNI execution; disabling hardware intrinsics explicitly skips all 12 cases.

The exact repro remains 70 bytes / 14 instructions; unmasked VNNI and ordinary masked-add controls have unchanged code. The existing conservative RMW exclusion also prevents zero-mask fusion: a compile-time-zero-fallback probe grows from 63 to 69 bytes and 9 to 10 instructions. Local benchmark timings were unstable, so no throughput claim is made. VEX-only VNNI and newer VNNI Int8/Int16 hardware paths were not executable on this host.

Resolves #133753

Note

This PR description was drafted with GitHub Copilot.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 21, 2026 18:54
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Newer INT8/INT16 variants lack regression coverage, and zero-fallback masked fusion regresses unnecessarily.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 Medium severity

Open (2)
What changed in this PR

This PR fixes incorrect JIT containment of masked VNNI multiply-add operations by marking them as read-modify-write intrinsics.

Changes:

  • Updates RMW metadata for regular, saturating, and newer VNNI variants.
  • Adds 12 regression cases across vector sizes and operand forms.
File Description
src/​coreclr/​jit/​hwintrinsiclistxarch.h Marks VNNI intrinsic entries as RMW.
src/​tests/​JIT/​Regression_ro_2/​Runtime_133753.cs Adds masked-blend regression coverage.

Comment thread src/coreclr/jit/hwintrinsiclistxarch.h
Comment thread src/tests/JIT/Regression_ro_2/Runtime_133753.cs
@tannergooding
tannergooding merged commit f2ecb6b into dotnet:main Sep 22, 2026
143 of 146 checks passed
@tannergooding

Copy link
Copy Markdown
Member Author

/backport to release/11.0

Note

Backport requested through GitHub Copilot on behalf of @tannergooding.

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/11.0 (link to workflow run)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

JIT: (bug) AvxVnni.V512.MultiplyWideningAndAdd folded into BlendVariable loses the merge operand

3 participants