Fix VNNI blend merge operand preservation - #134364
Merged
tannergooding merged 1 commit intoSep 22, 2026
Merged
tannergooding merged 1 commit into
tannergooding merged 1 commit into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Azure Pipelines: Successfully started running 5 pipeline(s). 11 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Newer INT8/INT16 variants lack regression coverage, and zero-fallback masked fusion regresses unnecessarily.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 2
Open (2)
What changed in this PR
This PR fixes incorrect JIT containment of masked VNNI multiply-add operations by marking them as read-modify-write intrinsics.
Changes:
- Updates RMW metadata for regular, saturating, and newer VNNI variants.
- Adds 12 regression cases across vector sizes and operand forms.
| File | Description |
|---|---|
src/coreclr/jit/hwintrinsiclistxarch.h |
Marks VNNI intrinsic entries as RMW. |
src/tests/JIT/Regression_ro_2/Runtime_133753.cs |
Adds masked-blend regression coverage. |
EgorBo
approved these changes
Sep 21, 2026
Member
Author
|
/backport to release/11.0 Note Backport requested through GitHub Copilot on behalf of @tannergooding. |
Contributor
|
Started backporting to |
4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

VNNI multiply-and-accumulate instructions read their destination as the accumulator. When merge-masked, inactive lanes also retain that destination. Folding
BlendVariable(fallback, MultiplyWideningAndAdd(addend, left, right), mask)into a masked VNNI instruction therefore cannot preserve an independent fallback: initializing the accumulator overwrites it.Mark the regular and saturating VNNI intrinsic entries as
HW_Flag_RmwIntrinsic, including the newer integer variants. This lets the existing embedded-masking compatibility check reject the invalid containment and emit the VNNI operation followed by the blend. No special-case lowering or codegen is needed.Add 12 regression cases covering 128/256/512-bit vectors, byte/short products, and saturating/non-saturating forms. Distinct fallback/addend values, zero products, and alternating comparison lanes isolate merge preservation. The repro helpers force optimized compilation even under default tiering.
Validation on Windows x64 with AVX-512 VNNI hardware:
The exact repro remains 70 bytes / 14 instructions; unmasked VNNI and ordinary masked-add controls have unchanged code. The existing conservative RMW exclusion also prevents zero-mask fusion: a compile-time-zero-fallback probe grows from 63 to 69 bytes and 9 to 10 instructions. Local benchmark timings were unstable, so no throughput claim is made. VEX-only VNNI and newer VNNI Int8/Int16 hardware paths were not executable on this host.
Resolves #133753
Note
This PR description was drafted with GitHub Copilot.