Skip to content

Add an opt-in ZeRO-1/2 gradient norm fast path - #8331

Merged
tohtana merged 4 commits into
deepspeedai:masterfrom
yh0903:yh0903/zero1-optimizer-fastpath
Sep 10, 2026
Merged

tohtana merged 4 commits into
deepspeedai:masterfrom
yh0903:yh0903/zero1-optimizer-fastpath

Conversation

@yh0903

@yh0903 yh0903 commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • add zero_optimization.compute_grad_norm, defaulting to true
  • allow ZeRO Stage 1/2 GPU optimizers to skip an unused global gradient norm while preserving overflow detection
  • skip the gradient multiply when clipping is disabled and the effective loss scale is exactly 1
  • return None from get_global_grad_norm() when norm computation is explicitly disabled
  • fail fast for gradient clipping, optimizer offload, ZenFlow, unsupported ZeRO stages, checkpoint-restored clipping, and the dedicated ZeRO-1 BF16 optimizer with FP32 gradient accumulation

Safety contract

Finite/overflow checking is unchanged. This does not remove the required non-finite scan; it only removes work with no mathematical consumer.

The default configuration and numerical behavior are unchanged. Identity unscaling is skipped automatically when clipping is disabled and the effective loss scale is exactly 1. Gradient-norm computation remains enabled by default because its cached result is externally observable through get_global_grad_norm().

Performance

Qwen3-30B-A3B, 48 layers, EP16 on 2x8 H100, sequence length 1024, BF16, ZeRO-1, activation checkpointing enabled:

Metric Baseline Fast path Change
Median step 987.92 ms 966.67 ms -2.2%
Optimizer update 84.93 ms 59.82 ms -29.6%
Global norm 20.30 ms 0 ms removed
Identity unscale 5.09 ms 0.03 ms removed

The paired block used a benchmark-only control that restores the pre-change multiply-by-one path. Peak allocated and reserved memory were unchanged. Two additional rotated blocks isolating the opt-in norm change were both positive (+3.3% and +0.6% E2E); their optimizer-phase saving was stable at about 20 ms, while full-step variance remained larger.

Testing Done

  • rebased onto current master
  • H100 targeted suite: 20 passed, covering Stage 1/2 update parity, overflow skip behavior, get_global_grad_norm(), clipping/offload/ZenFlow guards, checkpoint clipping, identity/non-identity scaling, and the ZeRO-1 BF16 optimizer with FP32 gradient accumulation
  • changed-file pre-commit hooks passed; flake8 5.0.4 was run separately under Python 3.13 because its pyflakes dependency is incompatible with Python 3.14
  • current-head multi-GPU H100 suite: 20 passed, 0 failed, 0 skipped
  • full repository CI still requires maintainer approval to run on the fork PR

yh0903 and others added 2 commits September 3, 2026 01:10
Preserve overflow detection while allowing ZeRO-1/2 GPU optimizers to skip an unused global norm, and avoid identity gradient scaling when clipping is disabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Fail fast when ZeRO-1 selects the dedicated BF16 optimizer with FP32 gradient accumulation, where disabling norm computation is not implemented.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
@yh0903
yh0903 force-pushed the yh0903/zero1-optimizer-fastpath branch from 86ae2dd to 349e8c9 Compare September 3, 2026 08:33
@yh0903
yh0903 marked this pull request as ready for review September 3, 2026 08:34

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 349e8c91fe

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread deepspeed/runtime/zero/config.py
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-03T08:37:30.299076Z 349e8c9 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Document that the dedicated ZeRO-1 BF16 optimizer is excluded when gradient norm computation is disabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>

@tohtana tohtana left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @yh0903,
Thank you for the improvement! Sorry for the delay of review.
This looks good to me.

@tohtana
tohtana enabled auto-merge September 9, 2026 01:10
@tohtana
tohtana added this pull request to the merge queue Sep 9, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 9, 2026
@tohtana
tohtana added this pull request to the merge queue Sep 9, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 9, 2026
@tohtana
tohtana added this pull request to the merge queue Sep 10, 2026
Merged via the queue into deepspeedai:master with commit f992dc2 Sep 10, 2026
13 checks passed
pull Bot pushed a commit to Abaso007/DeepSpeed that referenced this pull request Sep 22, 2026
## Summary

- add an opt-in `compile.autoep_non_moe` configuration option;
`engine.compile()` then regionally compiles callable parents of AutoEP
layers
- keep `AutoEPMoELayer.forward` as an explicit compiler-disabled graph
break, so routing, token movement, expert compute, and AllToAll
collectives remain eager
- discover regions from the injected AutoEP module hierarchy, including
a callable model root with a direct AutoEP child
- fail fast for unsupported configurations instead of silently changing
compile behavior
- accept explicitly disabled offload configurations while continuing to
reject active offload
- validate the production eager MoE boundary directly and compare actual
FP32 master-parameter updates

The option defaults to `false`, preserving the existing full-model
behavior of `engine.compile()`. Setting the option alone does not
trigger compilation. Enable it in the DeepSpeed configuration and call
`engine.compile()` after initialization:

```json
{
  "compile": {
    "autoep_non_moe": true
  }
}
```

```python
engine, optimizer, _, _ = deepspeed.initialize(model=model, config=ds_config)
engine.compile()
```

The decoder blocks contain the repeated attention, normalization,
residual and dense work targeted by this optimization. Selecting these
regions bounds the traced code and allows similar blocks to reuse
compiled graphs. Embeddings and the language-model head outside those
regions remain eager; compiling them would need separate validation and
performance measurements. A selected model-root region includes its own
non-MoE operations.

## Support boundary

This first version uses vanilla `torch.compile`/Inductor with the
standard AutoEP `comm` backend, sequence and pipeline parallel sizes of
one, and ZeRO stages 0, 1, and 2. Distributed performance and parity
validation currently target ZeRO stage 1.

DeepEP, DeepCompile, AutoEP+AutoTP folding, sequence or pipeline
parallelism, ZeRO stage 3, optimizer or parameter offload, compiled
autograd, DeepCompile schedules, and any `fullgraph` or `dynamic` value
other than `False` are rejected.

## Historical performance

These measurements predate this review follow-up and the latest master
merge; they are not a new performance measurement of the updated
revision.

Qwen3-30B-A3B, 48 layers, EP16, sequence length 1024, BF16, activation
checkpointing enabled, fixed routing, 2x8 H100. Each allocation
discarded one fixed warm arm and used two runs per variant.

| Allocation | Implementation / order | Eager median | Compiled median |
Throughput gain | Eager p95 median | Compiled p95 median |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| 1 | pre-production canary, ECCE | 936.66 ms | 885.82 ms | +5.74% |
1627.42 ms | 1534.37 ms |
| 2 | production engine API, CEEC | 932.61 ms | 875.10 ms | +6.57% |
1537.12 ms | 1446.36 ms |

Both allocations reduced peak allocated memory by 1.60 GiB and peak
reserved memory by 2.01 GiB. The maximum paired loss differences were
0.00617 and 0.00194. All 16 ranks reported zero Dynamo counter changes
in the measured window and exactly 2,880 eager AutoEP calls per compiled
arm (`30 steps x 48 layers x forward/replay`).

An exact same-stack Nsight census explained the clean E2E improvement:

| Metric | Eager | Compiled | Delta |
| --- | ---: | ---: | ---: |
| GPU kernels | 22,340 | 15,236 | -7,104 (-31.8%) |
| NCCL kernels | 390 | 390 | unchanged |
| non-NCCL kernel union | 233.73 ms | 191.42 ms | -42.31 ms |
| launch API time | 111.76 ms | 83.07 ms | -28.69 ms |

Compiled execution also removed 1.59 GiB of D2D traffic per captured
step. These performance runs used the PR1+PR2+PR3 benchmark stack
(deepspeedai#8326, deepspeedai#8331, and deepspeedai#8359) to isolate the remaining Transformer/runtime
fragmentation. This PR is based directly on `master` and has no code
dependency on those changes.

## Project rebaseline

This is context for the broader AutoEP optimization effort, not a merge
gate for this PR. A same-allocation fixed-routing `discard + ABBA +
BAAB` comparison of the full deepspeedai#8326 + deepspeedai#8331 + deepspeedai#8359 + deepspeedai#8380 stack against
Megatron Core measured:

| Framework | Median step | Median of per-arm p95 | Peak allocated |
Peak reserved |
| --- | ---: | ---: | ---: | ---: |
| AutoEP PR1-4 | 887.21 ms | 1244.49 ms | 40.37 GiB | 45.29 GiB |
| Megatron Core | 814.58 ms | 821.30 ms | 41.68 GiB | 45.80 GiB |

The residual in that historical comparison was **72.64 ms / AutoEP
+8.9%**. All four paired deltas favored Megatron and stayed within
**66.55–77.19 ms**; AutoEP used 1.30 GiB less allocated and 0.51 GiB
less reserved memory.

A separate low-overhead outer CUDA-event trace placed the typical
residual primarily in:

- forward: **AutoEP +39.14 ms**
- backward including ordinary gradient synchronization: **AutoEP +46.27
ms**
- optimizer: **AutoEP -9.00 ms**

The AutoEP tail is not a measured-window recompilation. In both traced
AutoEP arms, step 27's forward increased from a typical 266–271 ms to
678 ms while backward and optimizer remained near their medians. With
only ten measured steps per arm, the reported p95 is the maximum sample;
this deterministic forward-only spike is a follow-up investigation
rather than evidence that regional compile regressed the typical path.

## Validation

- 57 targeted CPU tests pass: configuration parsing, unchanged
full-model compile behavior, callable-root forward/backward parity with
checkpointing off/on, fail-fast and rollback contracts, parity
assertions, and existing DeepCompile lifecycle tests.
- Both callable-root regression cases reproduce the erroneous rejection
with the original helper and pass after the fix.
- `pre-commit run --files` passes for all seven changed files; all
non-merge commits carry `Signed-off-by` trailers.
- Exact-head GPU validation of
`85ba684c8709bfe1099c41151f504d83d7344636`: all **8/8 cases passed**,
with zero failures, errors or skips, on **2 × NVIDIA H100 80GB HBM3**
(PyTorch **2.10.0.2+cu130**, CUDA **13.0**, NCCL **2.28.9**). Coverage
is nested/root layouts × checkpointing off/on × async split planning
off/on, using BF16 and ZeRO stage 1.

The GPU suite uses the configuration option through
`deepspeed.initialize()` and the unchanged `engine.compile()` call. It
checks actual Inductor graph capture, output/loss/input gradients,
router/expert/non-MoE gradients, FP32 master-parameter updates, exact
routing assignments, the production eager AutoEP boundary, and no
additional Dynamo graphs/calls after warmup. These targeted tests are
correctness checks, not a new performance measurement. The single
measured SGD step uses `lr=1.0` to keep FP32 master updates above
rounding near unit-valued LayerNorm parameters. At `lr=0.01`, the
root-layout update norm was about `4e-7`, so a difference of one FP32
ULP exceeded the 5% relative gate. The forward/gradient comparisons and
every acceptance threshold are unchanged.

The current `modal-torch-latest` workflow runs test selection on PR
updates and executes its GPU suite on merge-queue entries. The
historical full-CI timeout is not counted as a completed test run.

---------

Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@tohtana tohtana mentioned this pull request Sep 27, 2026
10 of 22 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants