Conversation
Struct copies can write a promoted field back to its stack slot and immediately reload the same bits into another replacement. Copy equal-sized small integers with different signedness directly, relying on existing store normalization. Allow native-int/byref pairs while preserving each local's GC type, and skip partial-overlap write-backs when the source storage is already current. Object-reference reinterpretation remains excluded. Validation against this change's base on updated main a652cdf: - Windows x64 Checked and Release and Linux x64 Release JIT builds. - 18 existing physical-promotion regression runs pass with normal, forced promotion and JitStress=2 settings. - Focused optimized, promotion-stress, JitStress=2 and warmed-up tiering/PGO executions pass on base and candidate; Tier1 is captured. - Windows codegen checks fail on base and pass on candidate. Linux behavior passes, with codegen checks where applicable. - Byref GCStress=0xC runs pass on both JITs. - Eleven SuperPMI collections: 983,234 comparable contexts; 238,457,491 -> 238,457,491 bytes (0 saved). 0 shrink; 0 grow by 0 bytes total. Zero compilation failures; 190 missing-recording contexts excluded. The 190 recording gaps affect both sides. - Release PIN JIT instructions, benchmarks: +0.00008%. - Release PIN JIT instructions, HashSet PGO: -0.00067%. The corpus has no assembly differences; the focused FullOpts and Tier1 examples demonstrate the generated-code improvements. Code size, static PerfScore and JIT instruction counts do not establish application execution speed. No ARM64 or x86 execution is claimed. Related to dotnet#133833.
|
Azure Pipelines: Successfully started running 5 pipeline(s). 11 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
Contributor
There was a problem hiding this comment.
🔵 Needs a closer look
JIT promotion and GC/byref changes require final human review.
Pull request overview
Optimizes RyuJIT physical-promotion struct copies while preserving integer normalization and GC tracking.
Changes:
- Adds direct compatible field and native-int/byref copies.
- Skips redundant partial-overlap write-backs.
- Adds integer, byref, and GC-lifetime regression tests.
File summaries
| File | Description |
|---|---|
src/tests/JIT/Directed/physicalpromotion/CopyBetweenFields.csproj |
Configures field-copy tests. |
src/tests/JIT/Directed/physicalpromotion/CopyBetweenFields.cs |
Tests integer copies and write-back behavior. |
src/tests/JIT/Directed/physicalpromotion/CopyBetweenByrefs.csproj |
Configures byref tests. |
src/tests/JIT/Directed/physicalpromotion/CopyBetweenByrefs.cs |
Tests native-int/byref copying and GC safety. |
src/tests/JIT/Directed/Directed_do.csproj |
Registers the new test projects. |
src/coreclr/jit/promotiondecomposition.cpp |
Implements optimized replacement copying and conditional write-backs. |
Review details
- Files reviewed: 6/6 changed files
- Comments generated: 0
- Review effort level: Lite
Remove native-int/byref matching and the raw-byte reinterpretation tests. These tests do not establish a supported GC-tracking contract. Retain equal-sized small-integer signedness copies and redundant write-back elimination. Validation: Windows x64 Checked and Release and Linux x64 Release JIT builds; focused behavioral and codegen checks; 18 existing regression runs. All passed. The four retained examples still save 5-7 bytes in FullOpts and Tier1. SuperPMI: 983,234 comparable contexts, no assembly differences or compilation failures; 190 shared recording gaps. Release JIT instruction counts are effectively unchanged.
This was referenced Sep 14, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Struct copies can write a promoted field back to its stack slot and immediately reload the same bits into another replacement. Copy equal-sized small integers with different signedness directly, relying on existing store normalization. Skip partial-overlap write-backs when the source storage is already current.
Split from #133833 at @jakobbotsch's request so this change can be reviewed independently.
Base:
mainata652cdfed564.Focused generated code
Windows x64 method sizes. Tier1 measurements use a repeated harness with tiering and dynamic PGO enabled; the harness waits for tier-up and captures Tier1 output. FullOpts measurements disable tiering for deterministic codegen checks.
SignedByteToUnsignedSignedShortToUnsignedAtOffsetDirtySourceReadBackSourceAssembly example:
SignedByteToUnsignedFullOpts, Windows x64. Instruction bytes are omitted;
...denotes unchanged assembly.... mov ecx, ebx call [CopyBetweenFields:Observe(int):int] add esi, eax - mov byte ptr [rsp+0x20], bl - movzx rbx, byte ptr [rsp+0x20] + movzx rbx, bl mov ecx, ebx call [CopyBetweenFields:Observe(int):int] add esi, eax ... -; Total code size: 102 bytes +; Total code size: 96 bytesSuperPMI results
Windows x64 Checked JIT, eleven SuperPMI collections; identical settings on both sides and loop alignment disabled.
0 contexts shrink; 0 grow by 0 bytes total; 0 equal-size assembly differences. Zero compilation failures. The existing recordings contain 983,424 contexts. Updated main lacks recorded answers in 190 contexts on both sides; those contexts are excluded from the table.
The corpus has no assembly differences for this change. Its generated-code benefit is demonstrated by the focused methods above, including Tier1 compilation with tiering and dynamic PGO enabled.
Validation
Release JIT throughput was measured with PIN instruction counts (candidate versus this PR's base):
These small JIT instruction-count differences are not application-throughput measurements. Code size and static PerfScore also do not establish execution speed. No ARM64 or x86 execution is claimed.