Add an opt-in ZeRO-1/2 gradient norm fast path - #8331
Merged
tohtana merged 4 commits intoSep 10, 2026
Merged
Conversation
Preserve overflow detection while allowing ZeRO-1/2 GPU optimizers to skip an unused global norm, and avoid identity gradient scaling when clipping is disabled. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
Fail fast when ZeRO-1 selects the dedicated BF16 optimizer with FP32 gradient accumulation, where disabling norm computation is not implemented. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
yh0903
force-pushed
the
yh0903/zero1-optimizer-fastpath
branch
from
September 3, 2026 08:33
86ae2dd to
349e8c9
Compare
yh0903
marked this pull request as ready for review
September 3, 2026 08:34
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 349e8c91fe
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Document that the dedicated ZeRO-1 BF16 optimizer is excluded when gradient norm computation is disabled. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: yh0903 <helloyu0903@gmail.com>
tohtana
enabled auto-merge
September 9, 2026 01:10
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Sep 9, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Sep 9, 2026
8 tasks done
pull Bot
pushed a commit
to Abaso007/DeepSpeed
that referenced
this pull request
Sep 22, 2026
## Summary
- add an opt-in `compile.autoep_non_moe` configuration option;
`engine.compile()` then regionally compiles callable parents of AutoEP
layers
- keep `AutoEPMoELayer.forward` as an explicit compiler-disabled graph
break, so routing, token movement, expert compute, and AllToAll
collectives remain eager
- discover regions from the injected AutoEP module hierarchy, including
a callable model root with a direct AutoEP child
- fail fast for unsupported configurations instead of silently changing
compile behavior
- accept explicitly disabled offload configurations while continuing to
reject active offload
- validate the production eager MoE boundary directly and compare actual
FP32 master-parameter updates
The option defaults to `false`, preserving the existing full-model
behavior of `engine.compile()`. Setting the option alone does not
trigger compilation. Enable it in the DeepSpeed configuration and call
`engine.compile()` after initialization:
```json
{
"compile": {
"autoep_non_moe": true
}
}
```
```python
engine, optimizer, _, _ = deepspeed.initialize(model=model, config=ds_config)
engine.compile()
```
The decoder blocks contain the repeated attention, normalization,
residual and dense work targeted by this optimization. Selecting these
regions bounds the traced code and allows similar blocks to reuse
compiled graphs. Embeddings and the language-model head outside those
regions remain eager; compiling them would need separate validation and
performance measurements. A selected model-root region includes its own
non-MoE operations.
## Support boundary
This first version uses vanilla `torch.compile`/Inductor with the
standard AutoEP `comm` backend, sequence and pipeline parallel sizes of
one, and ZeRO stages 0, 1, and 2. Distributed performance and parity
validation currently target ZeRO stage 1.
DeepEP, DeepCompile, AutoEP+AutoTP folding, sequence or pipeline
parallelism, ZeRO stage 3, optimizer or parameter offload, compiled
autograd, DeepCompile schedules, and any `fullgraph` or `dynamic` value
other than `False` are rejected.
## Historical performance
These measurements predate this review follow-up and the latest master
merge; they are not a new performance measurement of the updated
revision.
Qwen3-30B-A3B, 48 layers, EP16, sequence length 1024, BF16, activation
checkpointing enabled, fixed routing, 2x8 H100. Each allocation
discarded one fixed warm arm and used two runs per variant.
| Allocation | Implementation / order | Eager median | Compiled median |
Throughput gain | Eager p95 median | Compiled p95 median |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| 1 | pre-production canary, ECCE | 936.66 ms | 885.82 ms | +5.74% |
1627.42 ms | 1534.37 ms |
| 2 | production engine API, CEEC | 932.61 ms | 875.10 ms | +6.57% |
1537.12 ms | 1446.36 ms |
Both allocations reduced peak allocated memory by 1.60 GiB and peak
reserved memory by 2.01 GiB. The maximum paired loss differences were
0.00617 and 0.00194. All 16 ranks reported zero Dynamo counter changes
in the measured window and exactly 2,880 eager AutoEP calls per compiled
arm (`30 steps x 48 layers x forward/replay`).
An exact same-stack Nsight census explained the clean E2E improvement:
| Metric | Eager | Compiled | Delta |
| --- | ---: | ---: | ---: |
| GPU kernels | 22,340 | 15,236 | -7,104 (-31.8%) |
| NCCL kernels | 390 | 390 | unchanged |
| non-NCCL kernel union | 233.73 ms | 191.42 ms | -42.31 ms |
| launch API time | 111.76 ms | 83.07 ms | -28.69 ms |
Compiled execution also removed 1.59 GiB of D2D traffic per captured
step. These performance runs used the PR1+PR2+PR3 benchmark stack
(deepspeedai#8326, deepspeedai#8331, and deepspeedai#8359) to isolate the remaining Transformer/runtime
fragmentation. This PR is based directly on `master` and has no code
dependency on those changes.
## Project rebaseline
This is context for the broader AutoEP optimization effort, not a merge
gate for this PR. A same-allocation fixed-routing `discard + ABBA +
BAAB` comparison of the full deepspeedai#8326 + deepspeedai#8331 + deepspeedai#8359 + deepspeedai#8380 stack against
Megatron Core measured:
| Framework | Median step | Median of per-arm p95 | Peak allocated |
Peak reserved |
| --- | ---: | ---: | ---: | ---: |
| AutoEP PR1-4 | 887.21 ms | 1244.49 ms | 40.37 GiB | 45.29 GiB |
| Megatron Core | 814.58 ms | 821.30 ms | 41.68 GiB | 45.80 GiB |
The residual in that historical comparison was **72.64 ms / AutoEP
+8.9%**. All four paired deltas favored Megatron and stayed within
**66.55–77.19 ms**; AutoEP used 1.30 GiB less allocated and 0.51 GiB
less reserved memory.
A separate low-overhead outer CUDA-event trace placed the typical
residual primarily in:
- forward: **AutoEP +39.14 ms**
- backward including ordinary gradient synchronization: **AutoEP +46.27
ms**
- optimizer: **AutoEP -9.00 ms**
The AutoEP tail is not a measured-window recompilation. In both traced
AutoEP arms, step 27's forward increased from a typical 266–271 ms to
678 ms while backward and optimizer remained near their medians. With
only ten measured steps per arm, the reported p95 is the maximum sample;
this deterministic forward-only spike is a follow-up investigation
rather than evidence that regional compile regressed the typical path.
## Validation
- 57 targeted CPU tests pass: configuration parsing, unchanged
full-model compile behavior, callable-root forward/backward parity with
checkpointing off/on, fail-fast and rollback contracts, parity
assertions, and existing DeepCompile lifecycle tests.
- Both callable-root regression cases reproduce the erroneous rejection
with the original helper and pass after the fix.
- `pre-commit run --files` passes for all seven changed files; all
non-merge commits carry `Signed-off-by` trailers.
- Exact-head GPU validation of
`85ba684c8709bfe1099c41151f504d83d7344636`: all **8/8 cases passed**,
with zero failures, errors or skips, on **2 × NVIDIA H100 80GB HBM3**
(PyTorch **2.10.0.2+cu130**, CUDA **13.0**, NCCL **2.28.9**). Coverage
is nested/root layouts × checkpointing off/on × async split planning
off/on, using BF16 and ZeRO stage 1.
The GPU suite uses the configuration option through
`deepspeed.initialize()` and the unchanged `engine.compile()` call. It
checks actual Inductor graph capture, output/loss/input gradients,
router/expert/non-MoE gradients, FP32 master-parameter updates, exact
routing assignments, the production eager AutoEP boundary, and no
additional Dynamo graphs/calls after warmup. These targeted tests are
correctness checks, not a new performance measurement. The single
measured SGD step uses `lr=1.0` to keep FP32 master updates above
rounding near unit-valued LayerNorm parameters. At `lr=0.01`, the
root-layout update norm was about `4e-7`, so a difference of one FP32
ULP exceeded the 5% relative gate. The forward/gradient comparisons and
every acceptance threshold are unchanged.
The current `modal-torch-latest` workflow runs test selection on PR
updates and executes its GPU suite on merge-queue entries. The
historical full-CI timeout is not counted as a completed test run.
---------
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
zero_optimization.compute_grad_norm, defaulting totrueNonefromget_global_grad_norm()when norm computation is explicitly disabledSafety contract
Finite/overflow checking is unchanged. This does not remove the required non-finite scan; it only removes work with no mathematical consumer.
The default configuration and numerical behavior are unchanged. Identity unscaling is skipped automatically when clipping is disabled and the effective loss scale is exactly 1. Gradient-norm computation remains enabled by default because its cached result is externally observable through
get_global_grad_norm().Performance
Qwen3-30B-A3B, 48 layers, EP16 on 2x8 H100, sequence length 1024, BF16, ZeRO-1, activation checkpointing enabled:
The paired block used a benchmark-only control that restores the pre-change multiply-by-one path. Peak allocated and reserved memory were unchanged. Two additional rotated blocks isolating the opt-in norm change were both positive (+3.3% and +0.6% E2E); their optimizer-phase saving was stable at about 20 ms, while full-step variance remained larger.
Testing Done
masterget_global_grad_norm(), clipping/offload/ZenFlow guards, checkpoint clipping, identity/non-identity scaling, and the ZeRO-1 BF16 optimizer with FP32 gradient accumulation