[release/11.0] Fix VNNI blend merge operand preservation - #134420
Open
github-actions[bot] wants to merge 1 commit into
Open
github-actions[bot] wants to merge 1 commit into
github-actions[bot] wants to merge 1 commit into
Conversation
VNNI multiply-and-accumulate instructions read their destination as the accumulator. When merge-masked, inactive lanes also retain that destination. Folding `BlendVariable(fallback, MultiplyWideningAndAdd(addend, left, right), mask)` into a masked VNNI instruction therefore cannot preserve an independent fallback: initializing the accumulator overwrites it. Mark the regular and saturating VNNI intrinsic entries as `HW_Flag_RmwIntrinsic`, including the newer integer variants. This lets the existing embedded-masking compatibility check reject the invalid containment and emit the VNNI operation followed by the blend. No special-case lowering or codegen is needed. Add 12 regression cases covering 128/256/512-bit vectors, byte/short products, and saturating/non-saturating forms. Distinct fallback/addend values, zero products, and alternating comparison lanes isolate merge preservation. The repro helpers force optimized compilation even under default tiering. Validation on Windows x64 with AVX-512 VNNI hardware: - Built Checked and Release JITs and the regression runner; ran JIT formatting. - All 12 cases fail with the original Checked JIT and pass with the fixed Checked JIT under default tiering. All 12 also pass with fixed Checked and Release JITs with tiering disabled. - Disassembly confirms actual VNNI execution; disabling hardware intrinsics explicitly skips all 12 cases. The exact repro remains 70 bytes / 14 instructions; unmasked VNNI and ordinary masked-add controls have unchanged code. The existing conservative RMW exclusion also prevents zero-mask fusion: a compile-time-zero-fallback probe grows from 63 to 69 bytes and 9 to 10 instructions. Local benchmark timings were unstable, so no throughput claim is made. VEX-only VNNI and newer VNNI Int8/Int16 hardware paths were not executable on this host. Resolves #133753 > [!NOTE] > This PR description was drafted with GitHub Copilot. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
EgorBo
approved these changes
Sep 22, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Backport of #134364 to
release/11.0.Marks VNNI multiply-and-accumulate intrinsics as read-modify-write so the existing embedded-masking compatibility check rejects invalid blend folding. Includes the same 12 regression cases as the main PR.
Customer Impact
Reported in #133753 against SDK
11.0.100-rc.2.26467.112. Optimized code that blends a VNNI multiply-add result with an independent fallback can silently return incorrect values on AVX-512-capable hardware. The folded instruction uses its destination for both the accumulator and inactive-lane merge value, losing the fallback. With fallback 7, addend 5, zero products, and alternating mask lanes, the expected result alternates 7 and 5; the affected code returns all 5. Both regular and saturating forms are affected. No external customer report is cited in the issue.Regression
The exact
AvxVnni.V512repro became possible with #128365 (8207ad3ca9a), which introduced that API in .NET 11 and enabled the existing narrower APIs on AVX-512-VNNI-only hardware.The underlying missing RMW classification predates .NET 11. The blend-folding machinery was introduced by #116983 (
96977c9353a) and is present in .NET 10. Source inspection indicates analogous 128/256-bit cases can reach the same faulty path there on hardware supporting both AVX-VNNI and AVX-512 VL. That .NET 10 assessment has not been execution-verified; this is not being classified as exclusively a .NET 11 regression.Testing
Validation for the main-branch fix on Windows x64 with an AMD Ryzen 9 7950X:
The previous coverage did not detect this particular blend-containment interaction. The backport bot applied the same two-file change without conflicts, and its diff was inspected. The release-branch backport has not been built or run locally; the results above are from the main-branch fix. VEX-only VNNI and newer VNNI Int8/Int16 hardware paths were not executable on the local host.
Risk
Low. The product change adds the existing
HW_Flag_RmwIntrinsicflag to eight intrinsic entries, using the existing RMW containment restriction rather than introducing new lowering or codegen logic. Regression cases demonstrate the original failure and the corrected results. Newer integer variants already have special codegen and are excluded from this folding path independently; their metadata is corrected consistently.There is a bounded optimization tradeoff: the existing RMW restriction also excludes zero-fallback fusion. A measured probe grows from 63 to 69 bytes and 9 to 10 instructions. Supporting that fusion safely requires separate register-allocation work. The exact reported repro remains 70 bytes / 14 instructions, and unmasked VNNI and ordinary masked-add controls have unchanged code. Local benchmark timings were unstable, so no throughput claim is made.
IMPORTANT: If this backport is for a servicing release, please verify that:
release/X.0-staging, notrelease/X.0.release/X.0(no-stagingsuffix).Verified target:
dotnet/runtime:release/11.0.Package authoring no longer needed in .NET 9
IMPORTANT: Starting with .NET 9, you no longer need to edit a NuGet package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older versions.
Resolves #133753
Note
This backport description was drafted with GitHub Copilot.