Skip to content

[release/11.0] Fix VNNI blend merge operand preservation - #134420

Open
github-actions[bot] wants to merge 1 commit into
release/11.0from
backport/pr-134364-to-release/11.0
Open

github-actions[bot] wants to merge 1 commit into
release/11.0from
backport/pr-134364-to-release/11.0

Conversation

@github-actions

@github-actions github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Backport of #134364 to release/11.0.

Marks VNNI multiply-and-accumulate intrinsics as read-modify-write so the existing embedded-masking compatibility check rejects invalid blend folding. Includes the same 12 regression cases as the main PR.

Customer Impact

  • Customer reported
  • Found internally

Reported in #133753 against SDK 11.0.100-rc.2.26467.112. Optimized code that blends a VNNI multiply-add result with an independent fallback can silently return incorrect values on AVX-512-capable hardware. The folded instruction uses its destination for both the accumulator and inactive-lane merge value, losing the fallback. With fallback 7, addend 5, zero products, and alternating mask lanes, the expected result alternates 7 and 5; the affected code returns all 5. Both regular and saturating forms are affected. No external customer report is cited in the issue.

Regression

  • Yes
  • No

The exact AvxVnni.V512 repro became possible with #128365 (8207ad3ca9a), which introduced that API in .NET 11 and enabled the existing narrower APIs on AVX-512-VNNI-only hardware.

The underlying missing RMW classification predates .NET 11. The blend-folding machinery was introduced by #116983 (96977c9353a) and is present in .NET 10. Source inspection indicates analogous 128/256-bit cases can reach the same faulty path there on hardware supporting both AVX-VNNI and AVX-512 VL. That .NET 10 assessment has not been execution-verified; this is not being classified as exclusively a .NET 11 regression.

Testing

Validation for the main-branch fix on Windows x64 with an AMD Ryzen 9 7950X:

  • Built Checked and Release JITs and the regression runner; JIT formatting passed.
  • All 12 final regression cases fail with the original Checked JIT and pass with the fixed Checked JIT under default tiering. All 12 also pass with fixed Checked and Release JITs with tiering disabled.
  • Cases cover 128/256/512-bit vectors, byte/short products, and saturating/non-saturating forms. Distinct fallback and accumulator values, zero products, and alternating masks isolate merge preservation. Optimized non-inlined helpers ensure the failing optimization is exercised.
  • Disassembly confirms actual AVX-512 VNNI execution. Disabling hardware intrinsics explicitly skips all 12 cases rather than counting unsupported execution as passing.

The previous coverage did not detect this particular blend-containment interaction. The backport bot applied the same two-file change without conflicts, and its diff was inspected. The release-branch backport has not been built or run locally; the results above are from the main-branch fix. VEX-only VNNI and newer VNNI Int8/Int16 hardware paths were not executable on the local host.

Risk

Low. The product change adds the existing HW_Flag_RmwIntrinsic flag to eight intrinsic entries, using the existing RMW containment restriction rather than introducing new lowering or codegen logic. Regression cases demonstrate the original failure and the corrected results. Newer integer variants already have special codegen and are excluded from this folding path independently; their metadata is corrected consistently.

There is a bounded optimization tradeoff: the existing RMW restriction also excludes zero-fallback fusion. A measured probe grows from 63 to 69 bytes and 9 to 10 instructions. Supporting that fusion safely requires separate register-allocation work. The exact reported repro remains 70 bytes / 14 instructions, and unmasked VNNI and ordinary masked-add controls have unchanged code. Local benchmark timings were unstable, so no throughput claim is made.

IMPORTANT: If this backport is for a servicing release, please verify that:

  • For .NET 8 and .NET 9: The PR target branch is release/X.0-staging, not release/X.0.
  • For .NET 10+: The PR target branch is release/X.0 (no -staging suffix).

Verified target: dotnet/runtime:release/11.0.

Package authoring no longer needed in .NET 9

IMPORTANT: Starting with .NET 9, you no longer need to edit a NuGet package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older versions.

Resolves #133753

Note

This backport description was drafted with GitHub Copilot.

VNNI multiply-and-accumulate instructions read their destination as the
accumulator. When merge-masked, inactive lanes also retain that
destination. Folding `BlendVariable(fallback,
MultiplyWideningAndAdd(addend, left, right), mask)` into a masked VNNI
instruction therefore cannot preserve an independent fallback:
initializing the accumulator overwrites it.

Mark the regular and saturating VNNI intrinsic entries as
`HW_Flag_RmwIntrinsic`, including the newer integer variants. This lets
the existing embedded-masking compatibility check reject the invalid
containment and emit the VNNI operation followed by the blend. No
special-case lowering or codegen is needed.

Add 12 regression cases covering 128/256/512-bit vectors, byte/short
products, and saturating/non-saturating forms. Distinct fallback/addend
values, zero products, and alternating comparison lanes isolate merge
preservation. The repro helpers force optimized compilation even under
default tiering.

Validation on Windows x64 with AVX-512 VNNI hardware:
- Built Checked and Release JITs and the regression runner; ran JIT
formatting.
- All 12 cases fail with the original Checked JIT and pass with the
fixed Checked JIT under default tiering. All 12 also pass with fixed
Checked and Release JITs with tiering disabled.
- Disassembly confirms actual VNNI execution; disabling hardware
intrinsics explicitly skips all 12 cases.

The exact repro remains 70 bytes / 14 instructions; unmasked VNNI and
ordinary masked-add controls have unchanged code. The existing
conservative RMW exclusion also prevents zero-mask fusion: a
compile-time-zero-fallback probe grows from 63 to 69 bytes and 9 to 10
instructions. Local benchmark timings were unstable, so no throughput
claim is made. VEX-only VNNI and newer VNNI Int8/Int16 hardware paths
were not executable on this host.

Resolves #133753

> [!NOTE]
> This PR description was drafted with GitHub Copilot.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 22, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants