You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
replace ReadyToRun's persistent raw-thread work list with Parallel.ForEach, relying on TPL to join active compilation work before returning or propagating an exception
use a no-buffering dynamic partitioner in both ReadyToRun and NativeAOT because method compilation costs vary widely and buffered partitions can cause high-core load imbalance
keep ReadyToRun's CorInfoImpl and recycling counter in partition-local state, bounding cache lifetime without tying large caches to ThreadPool threads
retain one ReadyToRun WorkerState across single-threaded compilation waves because SuperPMI collection uses --parallelism:1 and requires stable ObjectToHandle handles
dispose the ReadyToRun compilation when Compile throws and release single-threaded JIT state before object emission
Testing
./build.sh clr.jit+clr.hosts: succeeded
ILCompiler.ReadyToRun and ILCompiler.RyuJit builds: succeeded with 0 warnings and 0 errors
ReadyToRun compiler output was byte-for-byte identical between the dedicated-thread and TPL implementations
ReadyToRun CoreLib compilation performance was neutral: parallelism 4 median 22.26s vs. 22.76s; parallelism 24 median 10.03s vs. 9.94s
NativeAOT compiling crossgen2 with ILC (five alternating samples after warmup on a 32-logical-core Intel Xeon Platinum 8370C): parallelism 4 median 58.46s -> 57.09s (2.3% faster); parallelism 24 median 41.49s -> 39.91s (3.8% faster)
NativeAOT crossgen2 object output was byte-for-byte identical across both implementations and both parallelism settings; median peak RSS was comparable (p4: 1282 MiB -> 1244 MiB, p24: 1296 MiB -> 1316 MiB)
Note
This pull request description was generated with GitHub Copilot.
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
This PR refactors parallel ReadyToRun compilation to centralize work distribution/worker lifetime management in a new CompilationWorklist<TWorkerState>, improving exception handling and cleanup behavior, and adjusts crossgen2 compilation ownership to ensure deterministic disposal.
Changes:
Introduces CompilationWorklist<TWorkerState> to coordinate parallel work and propagate the first unexpected worker exception to the coordinating thread.
Refactors ReadyToRunCodegenCompilation to use per-worker state and dispose worker/JIT state earlier to reduce peak memory during object emission.
Refactors crossgen2 to build compilations via a helper and use scoped (using var) disposal for safer cleanup.
Run(...) should validate that the provided items list is non-null before storing it and using items.Count in TryTake; otherwise a null argument would crash later with a NullReferenceException.
public void Run(IReadOnlyList<DependencyNodeCore<NodeFactory>> items)
{
EnsureWorkers();
Debug.Assert(_items is null);
_items = items;
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
It changes core parallel compilation orchestration and error propagation behavior in crossgen2, which merits careful human validation of concurrency/exception semantics.
Review tier: Lite Findings: 1
New issues introduced by this change (1)
Severity
Finding
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs — The PR description mentions encapsulating work distribution/persistent worker lifetime in…
Issues resolved since last review (1)
Severity
Finding
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/CompilationWorklist.cs — Validate constructor arguments so invalid parallelism values (0/negative) and a null delegate fail… View resolved comment
Suppressed comments (2)
Previously missed (2) — in code that hasn't changed since the last review.
src/coreclr/tools/aot/crossgen2/Program.cs:356
determinismCheckFailed is only used to cache compilation.DeterminismCheckFailed and can be removed to simplify control flow without changing behavior. src/coreclr/tools/aot/crossgen2/Program.cs:378
BuildCompilation always builds a ReadyToRunCodegenCompilation (via ReadyToRunCodegenCompilationBuilder.ToCompilation()), so returning ICompilation here hides the concrete type and forces casts at call sites.
One or more custom setup steps configured for this repository failed during this Copilot code review run:
Install Dependencies
Setup steps run before each review. If the review above is missing context, or no review was posted at all, the failing step above may be the cause. See the workflow run for failure details, fix your setup steps configuration, and re-request a review.
Note
You can configure setup steps for Copilot code review separately from Copilot cloud agent with a copilot-code-review.yml file. Read the docs for details.
Replace the persistent custom compilation threads with Parallel.ForEach so each compilation wave is joined before errors propagate. Keep partition-local CorInfoImpl state for bounded cache lifetime and persistent single-thread state for SuperPMI handle stability. Ensure the compilation is disposed when compilation fails.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
Use the same no-buffering dynamic partitioner as ReadyToRun so compilation work with uneven method costs is distributed consistently between the two AOT compilers.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
It changes core parallel compilation control flow and exception propagation behavior in a high-blast-radius area that warrants focused human review.
Review tier: Lite Findings: 1 · 1
New issues introduced by this change (1)
Severity
Finding
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs — Parallel.ForEach will surface unexpected worker exceptions as an AggregateException, which can…
Pre-existing issues (1)
Severity
Finding
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs — The PR description mentions encapsulating work distribution/persistent worker lifetime in… View comment
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
There is a confirmed cross-thread flag updated with Volatile APIs but reset with a non-volatile write, which should be made consistent to avoid mixed volatile/non-volatile access.
Review tier: Lite Findings: 1 · 1
Pre-existing issues (2)
Severity
Finding
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs — Parallel.ForEach will surface unexpected worker exceptions as an AggregateException, which can… View comment
src/coreclr/tools/aot/ILCompiler.ReadyToRun/Compiler/ReadyToRunCodegenCompilation.cs — The PR description mentions encapsulating work distribution/persistent worker lifetime in… View comment
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
_compilationSessionGeneratedColdCode is written by worker threads via Volatile.Write and read via Volatile.Read; resetting it with a plain write introduces mixed volatile/non-volatile access on a cross-thread field. Reset it with Volatile.Write as well to keep the access pattern consistent and avoid reordering surprises.
This overload reports worker failures as an AggregateException. In DEBUG NativeAOT, ILCompilerRootCommand catches only CodeGenerationFailedException directly to print the single-method repro arguments, so an unresilient worker failure now bypasses that handler; the release path also gets only the generic aggregate message as its headline. Please unwrap and rethrow the original failure here or update the top-level handler to inspect InnerExceptions.
## Summary
- replace ReadyToRun's persistent raw-thread work list with
`Parallel.ForEach`, relying on TPL to join active compilation work
before returning or propagating an exception
- use a no-buffering dynamic partitioner in both ReadyToRun and
NativeAOT because method compilation costs vary widely and buffered
partitions can cause high-core load imbalance
- keep ReadyToRun's `CorInfoImpl` and recycling counter in
partition-local state, bounding cache lifetime without tying large
caches to ThreadPool threads
- retain one ReadyToRun `WorkerState` across single-threaded compilation
waves because SuperPMI collection uses `--parallelism:1` and requires
stable `ObjectToHandle` handles
- dispose the ReadyToRun compilation when `Compile` throws and release
single-threaded JIT state before object emission
## Testing
- `./build.sh clr.jit+clr.hosts`: succeeded
- `ILCompiler.ReadyToRun` and `ILCompiler.RyuJit` builds: succeeded with
0 warnings and 0 errors
- `ILCompiler.ReadyToRun.Tests`: 34 passed, 11 skipped, 0 failed
- `ILCompiler.Compiler.Tests`: 22 passed, 0 failed
- ReadyToRun compiler output was byte-for-byte identical between the
dedicated-thread and TPL implementations
- ReadyToRun CoreLib compilation performance was neutral: parallelism 4
median 22.26s vs. 22.76s; parallelism 24 median 10.03s vs. 9.94s
- NativeAOT compiling crossgen2 with ILC (five alternating samples after
warmup on a 32-logical-core Intel Xeon Platinum 8370C): parallelism 4
median 58.46s -> 57.09s (2.3% faster); parallelism 24 median 41.49s ->
39.91s (3.8% faster)
- NativeAOT crossgen2 object output was byte-for-byte identical across
both implementations and both parallelism settings; median peak RSS was
comparable (p4: 1282 MiB -> 1244 MiB, p24: 1296 MiB -> 1316 MiB)
> [!NOTE]
> This pull request description was generated with GitHub Copilot.
---------
Copilot-Session: b8b7894a-60fb-4c1a-9ace-1d5c5d6f6e01
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Parallel.ForEach, relying on TPL to join active compilation work before returning or propagating an exceptionCorInfoImpland recycling counter in partition-local state, bounding cache lifetime without tying large caches to ThreadPool threadsWorkerStateacross single-threaded compilation waves because SuperPMI collection uses--parallelism:1and requires stableObjectToHandlehandlesCompilethrows and release single-threaded JIT state before object emissionTesting
./build.sh clr.jit+clr.hosts: succeededILCompiler.ReadyToRunandILCompiler.RyuJitbuilds: succeeded with 0 warnings and 0 errorsILCompiler.ReadyToRun.Tests: 34 passed, 11 skipped, 0 failedILCompiler.Compiler.Tests: 22 passed, 0 failedNote
This pull request description was generated with GitHub Copilot.