Skip to content

JIT: Refine promotion costing for cold small parameter fields - #133603

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:fix-119770
Sep 16, 2026
Merged

jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:fix-119770

Conversation

@jakobbotsch

@jakobbotsch jakobbotsch commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

Withhold bitwise parameter-register extraction credit for byte/short fields whose weighted access count is less than 10% of the entry-block weight. Preserve direct register mappings and wider-field costing.

Many byte/short fields can fit into each parameter register, so eagerly extracting them can introduce substantial work and register pressure even when most fields are rarely accessed. Larger fields naturally limit the number of values that fit in a register and therefore the number of eager extractions. Restrict the heuristic to small fields to target this pattern.

This intentionally imprecise heuristic targets cases where eager field extraction is not profitable, such as Guid.CompareTo returning before its later fields are accessed. It does not precisely model extraction costs, register pressure, or shared spill/copy costs, and is not a general profitability model. Initialization of promoted fields remains unchanged.

Fix #119770

PriorityQueue<Guid, Guid> — fix-119770 vs baseline

Environment: Linux/x64, AMD Ryzen 9 5950X, Release JITs, tiered PGO enabled, ReadyToRun disabled. BenchmarkDotNet: two launches, eight warmups, fifteen measured iterations.

Benchmark Size Baseline fix-119770 Ratio BDN classification
HeapSort 10 186.2 ns 142.1 ns 0.76 Faster
Dequeue_And_Enqueue 10 559.2 ns 496.0 ns 0.89 Faster
K_Max_Elements 10 182.9 ns 158.3 ns 0.87 Faster
HeapSort 100 4.191 µs 2.785 µs 0.66 Faster
Dequeue_And_Enqueue 100 11.279 µs 9.234 µs 0.82 Faster
K_Max_Elements 100 846.0 ns 867.7 ns 1.03 Same
HeapSort 1000 83.540 µs 66.290 µs 0.79 Faster
Dequeue_And_Enqueue 1000 237.735 µs 212.093 µs 0.89 Faster
K_Max_Elements 1000 5.577 µs 4.592 µs 0.82 Faster

Withhold bitwise parameter-register extraction credit for byte/short fields
whose weighted access count is less than 10% of the entry-block weight.
Preserve direct register mappings and wider-field costing.

Many byte/short fields can fit into each parameter register, so eagerly
extracting them can introduce substantial work and register pressure even
when most fields are rarely accessed. Larger fields naturally limit the
number of values that fit in a register and therefore the number of eager
extractions. Restrict the heuristic to small fields to target this pattern.

This intentionally imprecise heuristic targets cases where eager field
extraction is not profitable, such as Guid.CompareTo returning before its
later fields are accessed. It does not precisely model extraction costs,
register pressure, or shared spill/copy costs, and is not a general
profitability model. Initialization of promoted fields remains unchanged.

Related to dotnet#119770.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 4c28b3c3-cf6c-4149-8eb1-964831492011
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 10, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment thread src/coreclr/jit/promotion.cpp Outdated
@jakobbotsch
jakobbotsch marked this pull request as ready for review September 15, 2026 09:05
Copilot AI lite review requested due to automatic review settings September 15, 2026 09:05
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@jakobbotsch

jakobbotsch commented Sep 15, 2026 •

Copy link
Copy Markdown
Member Author

cc @dotnet/jit-contrib PTAL @EgorBo

Minor diffs. This is really just targeting the regression we saw in benchmarks, which has to do with eagerly extracting lots of fields from Guid in Guid.CompareTo when PGO data shows that they won't be used (because the first check returns immediately).

I have a better fix in mind for .NET 12 where we do these extractions lazily. #133690 is the first step towards this.

@jakobbotsch
jakobbotsch requested a review from EgorBo September 15, 2026 09:09

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Add before/after evidence covering the reported regression and wider-field control cases.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Refines RyuJIT physical-promotion costing to avoid eagerly extracting rarely accessed small parameter fields.

Changes:

  • Adds extraction-credit control for parameter-register mapping.
  • Applies a 10% weighted-access threshold to cold byte/short fields.
  • Preserves direct mappings and wider-field behavior.
File summaries
File Summary
src/coreclr/jit/promotion.h Extends the parameter-register mapping declaration.
src/coreclr/jit/promotion.cpp Implements threshold-based extraction costing.
Review details
  • Files reviewed: 2/2 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread src/coreclr/jit/promotion.cpp
@jakobbotsch
jakobbotsch merged commit 0b942f4 into dotnet:main Sep 16, 2026
135 of 138 checks passed
@jakobbotsch

Copy link
Copy Markdown
Member Author

/backport to release/11.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/11.0 (link to workflow run)

JulieLeeMSFT pushed a commit that referenced this pull request Sep 16, 2026
… fields (#134033)

Backport of #133603 to release/11.0

/cc @jakobbotsch

## Customer Impact

- [ ] Customer reported
- [X] Found internally

Perf regression in code using `Guid` as keys.

## Regression

- [X] Yes
- [ ] No

Introduced in #133603 in .NET 11. This PR causes the JIT to eagerly
extract many fields from `Guid`, resulting in perf regressions in some
cases.

## Testing

Benchmarks in #133603 show improvements.

## Risk

Low. Adjust heuristic to avoid a lot of eager extractions.

Co-authored-by: Jakob Botsch Nielsen <Jakob.botsch.nielsen@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 4c28b3c3-cf6c-4149-8eb1-964831492011
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 17, 2026
jtschuster pushed a commit to jtschuster/runtime that referenced this pull request Sep 18, 2026
…#133603)

Withhold bitwise parameter-register extraction credit for byte/short
fields whose weighted access count is less than 10% of the entry-block
weight. Preserve direct register mappings and wider-field costing.

Many byte/short fields can fit into each parameter register, so eagerly
extracting them can introduce substantial work and register pressure
even when most fields are rarely accessed. Larger fields naturally limit
the number of values that fit in a register and therefore the number of
eager extractions. Restrict the heuristic to small fields to target this
pattern.

This intentionally imprecise heuristic targets cases where eager field
extraction is not profitable, such as Guid.CompareTo returning before
its later fields are accessed. It does not precisely model extraction
costs, register pressure, or shared spill/copy costs, and is not a
general profitability model. Initialization of promoted fields remains
unchanged.

Fix dotnet#119770

### PriorityQueue<Guid, Guid> — fix-119770 vs baseline

**Environment:** Linux/x64, AMD Ryzen 9 5950X, Release JITs, tiered PGO
enabled, ReadyToRun disabled. BenchmarkDotNet: two launches, eight
warmups, fifteen measured iterations.

| Benchmark | Size | Baseline | fix-119770 | Ratio | BDN classification
|
|---|---:|---:|---:|---:|---|
| HeapSort | 10 | 186.2 ns | 142.1 ns | 0.76 | Faster |
| Dequeue_And_Enqueue | 10 | 559.2 ns | 496.0 ns | 0.89 | Faster |
| K_Max_Elements | 10 | 182.9 ns | 158.3 ns | 0.87 | Faster |
| HeapSort | 100 | 4.191 µs | 2.785 µs | 0.66 | Faster |
| Dequeue_And_Enqueue | 100 | 11.279 µs | 9.234 µs | 0.82 | Faster |
| K_Max_Elements | 100 | 846.0 ns | 867.7 ns | 1.03 | Same |
| HeapSort | 1000 | 83.540 µs | 66.290 µs | 0.79 | Faster |
| Dequeue_And_Enqueue | 1000 | 237.735 µs | 212.093 µs | 0.89 | Faster |
| K_Max_Elements | 1000 | 5.577 µs | 4.592 µs | 0.82 | Faster |

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 4c28b3c3-cf6c-4149-8eb1-964831492011
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Perf] Linux/x64: 9 Regressions on 9/10/2025 11:38:35 AM +00:00

3 participants