JIT: Fix AVX state tracking for unrolled block operations - #134410
Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
|
Azure Pipelines: Successfully started running 6 pipeline(s). 10 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Two moderate findings remain: missing targeted regression coverage and overly broad AVX-width tracking for GC-containing initblk runs.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
What changed in this PR
Fixes x64 JIT AVX-state tracking for unrolled block operations.
Changes:
- Tracks SIMD width for unrolled memmove operations.
- Uses 128-bit zeroing for zero-initialized blocks.
- Updates AVX-state handling across the reviewed JIT code.
| File | Description |
|---|---|
src/coreclr/jit/lsraxarch.cpp |
Updates AVX-width tracking for block operations. |
src/coreclr/jit/codegenxarch.cpp |
Uses narrower zeroing instructions for initialization. |
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 19e82fc5-6aaa-4780-b66b-f08a2f5f9727
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Moderate findings remain regarding regression coverage and SIMD-width tracking for GC-containing initialization.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (1)
Resolved since last review (1)
|
PTAL @tannergooding @dotnet/jit-contrib diffs (new vzerouppers). Also, changed the SIMD zeroing to emit the 16b variant vs 32b to avoid some vzerouppers |
Follow up on #134410 by tracking AVX upper state from register-producing widths rather than vector types or memory-store widths. This removes unnecessary `vzeroupper` instructions while accounting for previously missed wide loads, copies, and integer-division intermediates. - Account for allocator-inserted spills/reloads and actual local register copies. - Narrow unsafe-widening/extraction copies to the defined source or required result width, and use XMM zeroing in integer-vector division. - Preserve the existing call/prolog/epilog clearing policy. Native-boundary policy changes, including the pre-existing unmanaged `calli` gap, remain out of scope. ### SuperPMI results [CI build 1610798](https://dev.azure.com/dnceng-public/public/_build/results?buildId=1610798), testing head `b20f2c3d058`: | Target | Compared contexts | Net code-size change | FullOpts contribution | |---|---:|---:|---:| | Windows x64 | 2,814,515 | +810 bytes | −1,008 bytes | | Linux x64 | 3,256,661 | +15,645 bytes | +11,970 bytes | | Windows x86 | 2,576,563 | +11,088 bytes | +8,784 bytes | ARM, ARM64, and WASM have no assembly diffs. Windows corpus coverage expanded substantially, so raw totals are not directly comparable with earlier runs. Remaining growth includes both corrected wide-producer accounting and conservative method-wide cleanup placement; it is not all individually necessary transition cleanup. JIT executed-instruction counts increased approximately 0.011–0.031% on xarch; these are compiler-work measurements, not application-throughput results. The assembly runs report zero failing compilations, with symmetric missing-context counts between base and diff. Validated with Checked x64/all-target JIT builds, focused execution and disassembly checks across default/AVX2/AVX/SSE/software configurations, targeted register-stress runs, and JIT formatting. Runtime execution was on Windows x64; cross-target replay is not runtime execution on those platforms. > [!NOTE] > This PR description and implementation were prepared with GitHub Copilot assistance. --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Fixes #133784