Skip to content

JIT: Fix AVX state tracking for unrolled block operations - #134410

Merged
EgorBo merged 4 commits into
dotnet:mainfrom
EgorBo:fix-jit-avx-state-133784
Sep 22, 2026
Merged

EgorBo merged 4 commits into
dotnet:mainfrom
EgorBo:fix-jit-avx-state-133784

Conversation

@EgorBo

@EgorBo EgorBo commented Sep 22, 2026

Copy link
Copy Markdown
Member

Fixes #133784

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
Copilot AI lite review requested due to automatic review settings September 22, 2026 11:13
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 22, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 6 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Two moderate findings remain: missing targeted regression coverage and overly broad AVX-width tracking for GC-containing initblk runs.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Fixes x64 JIT AVX-state tracking for unrolled block operations.

Changes:

  • Tracks SIMD width for unrolled memmove operations.
  • Uses 128-bit zeroing for zero-initialized blocks.
  • Updates AVX-state handling across the reviewed JIT code.
File Description
src/​coreclr/​jit/​lsraxarch.cpp Updates AVX-width tracking for block operations.
src/​coreclr/​jit/​codegenxarch.cpp Uses narrower zeroing instructions for initialization.

Comment thread src/coreclr/jit/lsraxarch.cpp
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
Copilot AI review requested due to automatic review settings September 22, 2026 11:36
EgorBo and others added 2 commits September 22, 2026 13:39
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Moderate findings remain regarding regression coverage and SIMD-width tracking for GC-containing initialization.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
Resolved since last review (1)

Comment thread src/coreclr/jit/lsraxarch.cpp
Copilot AI review requested due to automatic review settings September 22, 2026 11:45

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Add regression coverage and align SIMD-width accounting with per-run initblk code generation.

Review effort: Lite
Findings: None

Resolved since last review (1)

@EgorBo

EgorBo commented Sep 22, 2026

Copy link
Copy Markdown
Member Author

PTAL @tannergooding @dotnet/jit-contrib diffs (new vzerouppers).

Also, changed the SIMD zeroing to emit the 16b variant vs 32b to avoid some vzerouppers

@EgorBo
EgorBo merged commit d729c0a into dotnet:main Sep 22, 2026
142 of 143 checks passed
@EgorBo
EgorBo deleted the fix-jit-avx-state-133784 branch September 22, 2026 15:00
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 23, 2026
tannergooding added a commit that referenced this pull request Sep 25, 2026
Follow up on #134410 by tracking AVX upper state from
register-producing widths rather than vector types or memory-store
widths. This removes unnecessary `vzeroupper` instructions while
accounting for previously missed wide loads, copies, and
integer-division intermediates.

- Account for allocator-inserted spills/reloads and actual local
register copies.
- Narrow unsafe-widening/extraction copies to the defined source or
required result width, and use XMM zeroing in integer-vector division.
- Preserve the existing call/prolog/epilog clearing policy.
Native-boundary policy changes, including the pre-existing unmanaged
`calli` gap, remain out of scope.

### SuperPMI results

[CI build
1610798](https://dev.azure.com/dnceng-public/public/_build/results?buildId=1610798),
testing head `b20f2c3d058`:

| Target | Compared contexts | Net code-size change | FullOpts
contribution |
|---|---:|---:|---:|
| Windows x64 | 2,814,515 | +810 bytes | −1,008 bytes |
| Linux x64 | 3,256,661 | +15,645 bytes | +11,970 bytes |
| Windows x86 | 2,576,563 | +11,088 bytes | +8,784 bytes |

ARM, ARM64, and WASM have no assembly diffs. Windows corpus coverage
expanded substantially, so raw totals are not directly comparable with
earlier runs. Remaining growth includes both corrected wide-producer
accounting and conservative method-wide cleanup placement; it is not all
individually necessary transition cleanup.

JIT executed-instruction counts increased approximately 0.011–0.031% on
xarch; these are compiler-work measurements, not application-throughput
results. The assembly runs report zero failing compilations, with
symmetric missing-context counts between base and diff.

Validated with Checked x64/all-target JIT builds, focused execution and
disassembly checks across default/AVX2/AVX/SSE/software configurations,
targeted register-stress runs, and JIT formatting. Runtime execution was
on Windows x64; cross-target replay is not runtime execution on those
platforms.

> [!NOTE]
> This PR description and implementation were prepared with GitHub
Copilot assistance.

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

JIT: (bug) x64: unrolled memmove and initblk use YMM/ZMM but never emit vzeroupper

3 participants