Skip to content

Handle exceptions from parallel ReadyToRun compilation - #132282

Merged
agocke merged 4 commits into
dotnet:mainfrom
agocke:bail-out-on-exception
Sep 14, 2026
Merged

agocke merged 4 commits into
dotnet:mainfrom
agocke:bail-out-on-exception

Conversation

@agocke

@agocke agocke commented Aug 13, 2026 •

Copy link
Copy Markdown
Member

Summary

  • replace ReadyToRun's persistent raw-thread work list with Parallel.ForEach, relying on TPL to join active compilation work before returning or propagating an exception
  • use a no-buffering dynamic partitioner in both ReadyToRun and NativeAOT because method compilation costs vary widely and buffered partitions can cause high-core load imbalance
  • keep ReadyToRun's CorInfoImpl and recycling counter in partition-local state, bounding cache lifetime without tying large caches to ThreadPool threads
  • retain one ReadyToRun WorkerState across single-threaded compilation waves because SuperPMI collection uses --parallelism:1 and requires stable ObjectToHandle handles
  • dispose the ReadyToRun compilation when Compile throws and release single-threaded JIT state before object emission

Testing

  • ./build.sh clr.jit+clr.hosts: succeeded
  • ILCompiler.ReadyToRun and ILCompiler.RyuJit builds: succeeded with 0 warnings and 0 errors
  • ILCompiler.ReadyToRun.Tests: 34 passed, 11 skipped, 0 failed
  • ILCompiler.Compiler.Tests: 22 passed, 0 failed
  • ReadyToRun compiler output was byte-for-byte identical between the dedicated-thread and TPL implementations
  • ReadyToRun CoreLib compilation performance was neutral: parallelism 4 median 22.26s vs. 22.76s; parallelism 24 median 10.03s vs. 9.94s
  • NativeAOT compiling crossgen2 with ILC (five alternating samples after warmup on a 32-logical-core Intel Xeon Platinum 8370C): parallelism 4 median 58.46s -> 57.09s (2.3% faster); parallelism 24 median 41.49s -> 39.91s (3.8% faster)
  • NativeAOT crossgen2 object output was byte-for-byte identical across both implementations and both parallelism settings; median peak RSS was comparable (p4: 1282 MiB -> 1244 MiB, p24: 1296 MiB -> 1316 MiB)

Note

This pull request description was generated with GitHub Copilot.

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

Comment thread src/coreclr/tools/aot/crossgen2/Program.cs Outdated
@agocke
agocke marked this pull request as ready for review August 13, 2026 16:21

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR refactors parallel ReadyToRun compilation to centralize work distribution/worker lifetime management in a new CompilationWorklist<TWorkerState>, improving exception handling and cleanup behavior, and adjusts crossgen2 compilation ownership to ensure deterministic disposal.

Changes:

  • Introduces CompilationWorklist<TWorkerState> to coordinate parallel work and propagate the first unexpected worker exception to the coordinating thread.
  • Refactors ReadyToRunCodegenCompilation to use per-worker state and dispose worker/JIT state earlier to reduce peak memory during object emission.
  • Refactors crossgen2 to build compilations via a helper and use scoped (using var) disposal for safer cleanup.
Show a summary per file
File Description
src/coreclr/tools/aot/ILCompiler.ReadyToRun/ILCompiler.ReadyToRun.csproj Adds the new CompilationWorklist.cs to the build.
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs Replaces bespoke thread/semaphore orchestration with CompilationWorklist, adds per-worker state tracking, and disposes worklist before emission.
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/CompilationWorklist.cs New helper to manage parallel work distribution, worker lifetime, and exception capture/rethrow.
src/coreclr/tools/aot/crossgen2/Program.cs Refactors compilation creation into BuildCompilation and uses scoped disposal to simplify ownership and cleanup.

Review details

Suppressed comments (1)

src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/CompilationWorklist.cs:60

  • Run(...) should validate that the provided items list is non-null before storing it and using items.Count in TryTake; otherwise a null argument would crash later with a NullReferenceException.
        public void Run(IReadOnlyList<DependencyNodeCore<NodeFactory>> items)
        {
            EnsureWorkers();
            Debug.Assert(_items is null);

            _items = items;
  • Files reviewed: 4/4 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/CompilationWorklist.cs Outdated
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

Comment thread src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/CompilationWorklist.cs Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It changes core parallel compilation orchestration and error propagation behavior in crossgen2, which merits careful human validation of concurrency/exception semantics.

Review tier: Lite
Findings: 1 Low severity

New issues introduced by this change (1)
Severity Finding
Low severity src/​coreclr/​tools/​aot/​ILCompiler.ReadyToRun/​Compiler/​ReadyToRunCodegenCompilation.cs — The PR description mentions encapsulating work distribution/persistent worker lifetime in…
Issues resolved since last review (1)
Severity Finding
Medium severity src/​coreclr/​tools/​aot/​ILCompiler.ReadyToRun/​Compiler/​CompilationWorklist.cs — Validate constructor arguments so invalid parallelism values (0/negative) and a null delegate fail… View resolved comment
Suppressed comments (2)

Previously missed (2) — in code that hasn't changed since the last review.

src/coreclr/tools/aot/crossgen2/Program.cs:356

  • determinismCheckFailed is only used to cache compilation.DeterminismCheckFailed and can be removed to simplify control flow without changing behavior.
    src/coreclr/tools/aot/crossgen2/Program.cs:378
  • BuildCompilation always builds a ReadyToRunCodegenCompilation (via ReadyToRunCodegenCompilationBuilder.ToCompilation()), so returning ICompilation here hides the concrete type and forces casts at call sites.

Copilot AI commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

One or more custom setup steps configured for this repository failed during this Copilot code review run:

Install Dependencies

Setup steps run before each review. If the review above is missing context, or no review was posted at all, the failing step above may be the cause. See the workflow run for failure details, fix your setup steps configuration, and re-request a review.

Note

You can configure setup steps for Copilot code review separately from Copilot cloud agent with a copilot-code-review.yml file. Read the docs for details.

Replace the persistent custom compilation threads with Parallel.ForEach so each compilation wave is joined before errors propagate. Keep partition-local CorInfoImpl state for bounded cache lifetime and persistent single-thread state for SuperPMI handle stability. Ensure the compilation is disposed when compilation fails.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
@agocke
agocke force-pushed the bail-out-on-exception branch from e37a63b to c32e6dc Compare September 9, 2026 18:19
Copilot AI review requested due to automatic review settings September 9, 2026 18:19
Use the same no-buffering dynamic partitioner as ReadyToRun so compilation work with uneven method costs is distributed consistently between the two AOT compilers.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It changes core parallel compilation control flow and exception propagation behavior in a high-blast-radius area that warrants focused human review.

Review tier: Lite
Findings: 1 Medium severity · 1 Low severity

New issues introduced by this change (1)
Severity Finding
Medium severity src/​coreclr/​tools/​aot/​ILCompiler.ReadyToRun/​Compiler/​ReadyToRunCodegenCompilation.cs — Parallel.ForEach will surface unexpected worker exceptions as an AggregateException, which can…
Pre-existing issues (1)
Severity Finding
Low severity src/​coreclr/​tools/​aot/​ILCompiler.ReadyToRun/​Compiler/​ReadyToRunCodegenCompilation.cs — The PR description mentions encapsulating work distribution/persistent worker lifetime in… View comment

Copilot AI review requested due to automatic review settings September 9, 2026 18:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

There is a confirmed cross-thread flag updated with Volatile APIs but reset with a non-volatile write, which should be made consistent to avoid mixed volatile/non-volatile access.

Review tier: Lite
Findings: 1 Medium severity · 1 Low severity

Pre-existing issues (2)
Severity Finding
Medium severity src/​coreclr/​tools/​aot/​ILCompiler.ReadyToRun/​Compiler/​ReadyToRunCodegenCompilation.cs — Parallel.ForEach will surface unexpected worker exceptions as an AggregateException, which can… View comment
Low severity src/​coreclr/​tools/​aot/​ILCompiler.ReadyToRun/​Compiler/​ReadyToRunCodegenCompilation.cs — The PR description mentions encapsulating work distribution/persistent worker lifetime in… View comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs:869

  • _compilationSessionGeneratedColdCode is written by worker threads via Volatile.Write and read via Volatile.Read; resetting it with a plain write introduces mixed volatile/non-volatile access on a cross-thread field. Reset it with Volatile.Write as well to keep the access pattern consistent and avoid reordering surprises.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
Copilot AI review requested due to automatic review settings September 11, 2026 22:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

One or more issues must be addressed before approval.

Review tier: Lite
Findings: None

Resolved findings (2)
Previously missed findings (1)

In code that hasn't changed since last review

Medium severity Preserve code-generation failures across the parallel loop

src/​coreclr/​tools/​aot/​ILCompiler.RyuJit/​Compiler/​RyuJitCompilation.cs:176

This overload reports worker failures as an AggregateException. In DEBUG NativeAOT, ILCompilerRootCommand catches only CodeGenerationFailedException directly to print the single-method repro arguments, so an unresilient worker failure now bypasses that handler; the release path also gets only the generic aggregate message as its headline. Please unwrap and rethrow the original failure here or update the top-level handler to inspect InnerExceptions.

@jkotas jkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@agocke

agocke commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

/ba-g failure is #133798

@agocke
agocke merged commit 3c245fd into dotnet:main Sep 14, 2026
108 of 110 checks passed
@agocke
agocke deleted the bail-out-on-exception branch September 14, 2026 22:15
@github-project-automation github-project-automation Bot moved this to Done in AppModel Sep 14, 2026
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 15, 2026
jtschuster pushed a commit to jtschuster/runtime that referenced this pull request Sep 18, 2026
## Summary

- replace ReadyToRun's persistent raw-thread work list with
`Parallel.ForEach`, relying on TPL to join active compilation work
before returning or propagating an exception
- use a no-buffering dynamic partitioner in both ReadyToRun and
NativeAOT because method compilation costs vary widely and buffered
partitions can cause high-core load imbalance
- keep ReadyToRun's `CorInfoImpl` and recycling counter in
partition-local state, bounding cache lifetime without tying large
caches to ThreadPool threads
- retain one ReadyToRun `WorkerState` across single-threaded compilation
waves because SuperPMI collection uses `--parallelism:1` and requires
stable `ObjectToHandle` handles
- dispose the ReadyToRun compilation when `Compile` throws and release
single-threaded JIT state before object emission

## Testing

- `./build.sh clr.jit+clr.hosts`: succeeded
- `ILCompiler.ReadyToRun` and `ILCompiler.RyuJit` builds: succeeded with
0 warnings and 0 errors
- `ILCompiler.ReadyToRun.Tests`: 34 passed, 11 skipped, 0 failed
- `ILCompiler.Compiler.Tests`: 22 passed, 0 failed
- ReadyToRun compiler output was byte-for-byte identical between the
dedicated-thread and TPL implementations
- ReadyToRun CoreLib compilation performance was neutral: parallelism 4
median 22.26s vs. 22.76s; parallelism 24 median 10.03s vs. 9.94s
- NativeAOT compiling crossgen2 with ILC (five alternating samples after
warmup on a 32-logical-core Intel Xeon Platinum 8370C): parallelism 4
median 58.46s -> 57.09s (2.3% faster); parallelism 24 median 41.49s ->
39.91s (3.8% faster)
- NativeAOT crossgen2 object output was byte-for-byte identical across
both implementations and both parallelism settings; median peak RSS was
comparable (p4: 1282 MiB -> 1244 MiB, p24: 1296 MiB -> 1316 MiB)

> [!NOTE]
> This pull request description was generated with GitHub Copilot.

---------

Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants