Add opt-in regional compilation around AutoEP layers - #8380
Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
…moe-regional-compile
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
Accept explicitly disabled offload configs, observe the production eager boundary directly, and make small gradient and FP32 master-update parity checks discriminating. Signed-off-by: yh0903 <helloyu0903@gmail.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 04cc82b71c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| # DeepSpeed Team | ||
|
|
||
| from __future__ import annotations |
There was a problem hiding this comment.
Add the required Signed-off-by trailer
This is a non-merge commit, but its message contains no Signed-off-by: trailer. The repository requires every non-merge commit to be created with --signoff, so this commit does not satisfy the merge requirements; recreate it using the configured author name and email with --signoff.
AGENTS.md reference: AGENTS.md:L8-L8
Useful? React with 👍 / 👎.
tohtana
left a comment
There was a problem hiding this comment.
Thank you for the improvement, @yh0903!
It's really good to have an official support to combine MoE and compile in DeepSpeed.
The change overall looks good to me. I left one comment. Here are two points I'd like to discuss:
(1) In terms of the interface to enable this feature, adding compile_mode to compile seems a bit too much as it is currently only for this feature. How about adding an option to compile section in the config like:
{
"compile": {
"autoep_non_moe": true
}
}I think we could also make this the default behavior.
(2) This feature currently compiles only decoder blocks, while excluding other layers including embedding and lm_heads. Can you explain the intension?
| else: | ||
| module.forward = original_forward | ||
| for region, compiled_call in original_compiled_calls.items(): | ||
| region._compiled_call_impl = compiled_call |
There was a problem hiding this comment.
This also rejects a direct AutoEP child such as mlp, which is supported by custom AutoEP replacement. In that case, the separator is empty, but the parent is the callable model root (named_modules[""]). Only an empty module_name means the AutoEP layer itself is the root. Could we use if not module_name: here and add a regression test for a callable root with a direct AutoEP child?
Feel free to refer to yh0903#1 if it's useful.
There was a problem hiding this comment.
Thanks for the review! And sorry for the late response as I was working on the revision. I moved the opt-in into the DeepSpeed config as "compile": {"autoep_non_moe": true} and restored the existing engine.compile() signature. I kept the option disabled by default for this initial version so existing full-model compilation and configurations outside the supported AutoEP subset retain their current behavior.
I also fixed region discovery to allow a callable model root with a direct AutoEP child, following your suggestion. An AutoEP layer that is itself the model root remains rejected because it has no surrounding region. Regression coverage now exercises the root case through actual compilation and compares forward/backward results.
The decoder-region scope was intentional: repeated decoder blocks contain the attention, normalization, residual and dense work targeted by this optimization, while similar blocks can reuse compiled graphs. Embeddings and the lm_head are not inherently incompatible with compilation; they remain eager when outside the selected regions because extending the scope needs separate validation and measurement. A selected root region includes its non-MoE operations. I clarified this in the documentation.
The changes are in the latest commit with a follow-up test-fixture correction to keep single-step FP32 updates above rounding. All tolerances are unchanged. The final revision passes 57 targeted CPU tests and all eight Inductor parity cases on two H100s. Thanks!
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
tohtana
left a comment
There was a problem hiding this comment.
Thank you for the update, looks good to me.
Summary
compile.autoep_non_moeconfiguration option;engine.compile()then regionally compiles callable parents of AutoEP layersAutoEPMoELayer.forwardas an explicit compiler-disabled graph break, so routing, token movement, expert compute, and AllToAll collectives remain eagerThe option defaults to
false, preserving the existing full-model behavior ofengine.compile(). Setting the option alone does not trigger compilation. Enable it in the DeepSpeed configuration and callengine.compile()after initialization:{ "compile": { "autoep_non_moe": true } }The decoder blocks contain the repeated attention, normalization, residual and dense work targeted by this optimization. Selecting these regions bounds the traced code and allows similar blocks to reuse compiled graphs. Embeddings and the language-model head outside those regions remain eager; compiling them would need separate validation and performance measurements. A selected model-root region includes its own non-MoE operations.
Support boundary
This first version uses vanilla
torch.compile/Inductor with the standard AutoEPcommbackend, sequence and pipeline parallel sizes of one, and ZeRO stages 0, 1, and 2. Distributed performance and parity validation currently target ZeRO stage 1.DeepEP, DeepCompile, AutoEP+AutoTP folding, sequence or pipeline parallelism, ZeRO stage 3, optimizer or parameter offload, compiled autograd, DeepCompile schedules, and any
fullgraphordynamicvalue other thanFalseare rejected.Historical performance
These measurements predate this review follow-up and the latest master merge; they are not a new performance measurement of the updated revision.
Qwen3-30B-A3B, 48 layers, EP16, sequence length 1024, BF16, activation checkpointing enabled, fixed routing, 2x8 H100. Each allocation discarded one fixed warm arm and used two runs per variant.
Both allocations reduced peak allocated memory by 1.60 GiB and peak reserved memory by 2.01 GiB. The maximum paired loss differences were 0.00617 and 0.00194. All 16 ranks reported zero Dynamo counter changes in the measured window and exactly 2,880 eager AutoEP calls per compiled arm (
30 steps x 48 layers x forward/replay).An exact same-stack Nsight census explained the clean E2E improvement:
Compiled execution also removed 1.59 GiB of D2D traffic per captured step. These performance runs used the PR1+PR2+PR3 benchmark stack (#8326, #8331, and #8359) to isolate the remaining Transformer/runtime fragmentation. This PR is based directly on
masterand has no code dependency on those changes.Project rebaseline
This is context for the broader AutoEP optimization effort, not a merge gate for this PR. A same-allocation fixed-routing
discard + ABBA + BAABcomparison of the full #8326 + #8331 + #8359 + #8380 stack against Megatron Core measured:The residual in that historical comparison was 72.64 ms / AutoEP +8.9%. All four paired deltas favored Megatron and stayed within 66.55–77.19 ms; AutoEP used 1.30 GiB less allocated and 0.51 GiB less reserved memory.
A separate low-overhead outer CUDA-event trace placed the typical residual primarily in:
The AutoEP tail is not a measured-window recompilation. In both traced AutoEP arms, step 27's forward increased from a typical 266–271 ms to 678 ms while backward and optimizer remained near their medians. With only ten measured steps per arm, the reported p95 is the maximum sample; this deterministic forward-only spike is a follow-up investigation rather than evidence that regional compile regressed the typical path.
Validation
pre-commit run --filespasses for all seven changed files; all non-merge commits carrySigned-off-bytrailers.85ba684c8709bfe1099c41151f504d83d7344636: all 8/8 cases passed, with zero failures, errors or skips, on 2 × NVIDIA H100 80GB HBM3 (PyTorch 2.10.0.2+cu130, CUDA 13.0, NCCL 2.28.9). Coverage is nested/root layouts × checkpointing off/on × async split planning off/on, using BF16 and ZeRO stage 1.The GPU suite uses the configuration option through
deepspeed.initialize()and the unchangedengine.compile()call. It checks actual Inductor graph capture, output/loss/input gradients, router/expert/non-MoE gradients, FP32 master-parameter updates, exact routing assignments, the production eager AutoEP boundary, and no additional Dynamo graphs/calls after warmup. These targeted tests are correctness checks, not a new performance measurement. The single measured SGD step useslr=1.0to keep FP32 master updates above rounding near unit-valued LayerNorm parameters. Atlr=0.01, the root-layout update norm was about4e-7, so a difference of one FP32 ULP exceeded the 5% relative gate. The forward/gradient comparisons and every acceptance threshold are unchanged.The current
modal-torch-latestworkflow runs test selection on PR updates and executes its GPU suite on merge-queue entries. The historical full-CI timeout is not counted as a completed test run.