Skip to content

Add opt-in regional compilation around AutoEP layers - #8380

Merged
tohtana merged 10 commits into
deepspeedai:masterfrom
yh0903:yh0903/autoep-nonmoe-regional-compile
Sep 22, 2026
Merged

tohtana merged 10 commits into
deepspeedai:masterfrom
yh0903:yh0903/autoep-nonmoe-regional-compile

Conversation

@yh0903

@yh0903 yh0903 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • add an opt-in compile.autoep_non_moe configuration option; engine.compile() then regionally compiles callable parents of AutoEP layers
  • keep AutoEPMoELayer.forward as an explicit compiler-disabled graph break, so routing, token movement, expert compute, and AllToAll collectives remain eager
  • discover regions from the injected AutoEP module hierarchy, including a callable model root with a direct AutoEP child
  • fail fast for unsupported configurations instead of silently changing compile behavior
  • accept explicitly disabled offload configurations while continuing to reject active offload
  • validate the production eager MoE boundary directly and compare actual FP32 master-parameter updates

The option defaults to false, preserving the existing full-model behavior of engine.compile(). Setting the option alone does not trigger compilation. Enable it in the DeepSpeed configuration and call engine.compile() after initialization:

{
  "compile": {
    "autoep_non_moe": true
  }
}
engine, optimizer, _, _ = deepspeed.initialize(model=model, config=ds_config)
engine.compile()

The decoder blocks contain the repeated attention, normalization, residual and dense work targeted by this optimization. Selecting these regions bounds the traced code and allows similar blocks to reuse compiled graphs. Embeddings and the language-model head outside those regions remain eager; compiling them would need separate validation and performance measurements. A selected model-root region includes its own non-MoE operations.

Support boundary

This first version uses vanilla torch.compile/Inductor with the standard AutoEP comm backend, sequence and pipeline parallel sizes of one, and ZeRO stages 0, 1, and 2. Distributed performance and parity validation currently target ZeRO stage 1.

DeepEP, DeepCompile, AutoEP+AutoTP folding, sequence or pipeline parallelism, ZeRO stage 3, optimizer or parameter offload, compiled autograd, DeepCompile schedules, and any fullgraph or dynamic value other than False are rejected.

Historical performance

These measurements predate this review follow-up and the latest master merge; they are not a new performance measurement of the updated revision.

Qwen3-30B-A3B, 48 layers, EP16, sequence length 1024, BF16, activation checkpointing enabled, fixed routing, 2x8 H100. Each allocation discarded one fixed warm arm and used two runs per variant.

Allocation Implementation / order Eager median Compiled median Throughput gain Eager p95 median Compiled p95 median
1 pre-production canary, ECCE 936.66 ms 885.82 ms +5.74% 1627.42 ms 1534.37 ms
2 production engine API, CEEC 932.61 ms 875.10 ms +6.57% 1537.12 ms 1446.36 ms

Both allocations reduced peak allocated memory by 1.60 GiB and peak reserved memory by 2.01 GiB. The maximum paired loss differences were 0.00617 and 0.00194. All 16 ranks reported zero Dynamo counter changes in the measured window and exactly 2,880 eager AutoEP calls per compiled arm (30 steps x 48 layers x forward/replay).

An exact same-stack Nsight census explained the clean E2E improvement:

Metric Eager Compiled Delta
GPU kernels 22,340 15,236 -7,104 (-31.8%)
NCCL kernels 390 390 unchanged
non-NCCL kernel union 233.73 ms 191.42 ms -42.31 ms
launch API time 111.76 ms 83.07 ms -28.69 ms

Compiled execution also removed 1.59 GiB of D2D traffic per captured step. These performance runs used the PR1+PR2+PR3 benchmark stack (#8326, #8331, and #8359) to isolate the remaining Transformer/runtime fragmentation. This PR is based directly on master and has no code dependency on those changes.

Project rebaseline

This is context for the broader AutoEP optimization effort, not a merge gate for this PR. A same-allocation fixed-routing discard + ABBA + BAAB comparison of the full #8326 + #8331 + #8359 + #8380 stack against Megatron Core measured:

Framework Median step Median of per-arm p95 Peak allocated Peak reserved
AutoEP PR1-4 887.21 ms 1244.49 ms 40.37 GiB 45.29 GiB
Megatron Core 814.58 ms 821.30 ms 41.68 GiB 45.80 GiB

The residual in that historical comparison was 72.64 ms / AutoEP +8.9%. All four paired deltas favored Megatron and stayed within 66.55–77.19 ms; AutoEP used 1.30 GiB less allocated and 0.51 GiB less reserved memory.

A separate low-overhead outer CUDA-event trace placed the typical residual primarily in:

  • forward: AutoEP +39.14 ms
  • backward including ordinary gradient synchronization: AutoEP +46.27 ms
  • optimizer: AutoEP -9.00 ms

The AutoEP tail is not a measured-window recompilation. In both traced AutoEP arms, step 27's forward increased from a typical 266–271 ms to 678 ms while backward and optimizer remained near their medians. With only ten measured steps per arm, the reported p95 is the maximum sample; this deterministic forward-only spike is a follow-up investigation rather than evidence that regional compile regressed the typical path.

Validation

  • 57 targeted CPU tests pass: configuration parsing, unchanged full-model compile behavior, callable-root forward/backward parity with checkpointing off/on, fail-fast and rollback contracts, parity assertions, and existing DeepCompile lifecycle tests.
  • Both callable-root regression cases reproduce the erroneous rejection with the original helper and pass after the fix.
  • pre-commit run --files passes for all seven changed files; all non-merge commits carry Signed-off-by trailers.
  • Exact-head GPU validation of 85ba684c8709bfe1099c41151f504d83d7344636: all 8/8 cases passed, with zero failures, errors or skips, on 2 × NVIDIA H100 80GB HBM3 (PyTorch 2.10.0.2+cu130, CUDA 13.0, NCCL 2.28.9). Coverage is nested/root layouts × checkpointing off/on × async split planning off/on, using BF16 and ZeRO stage 1.

The GPU suite uses the configuration option through deepspeed.initialize() and the unchanged engine.compile() call. It checks actual Inductor graph capture, output/loss/input gradients, router/expert/non-MoE gradients, FP32 master-parameter updates, exact routing assignments, the production eager AutoEP boundary, and no additional Dynamo graphs/calls after warmup. These targeted tests are correctness checks, not a new performance measurement. The single measured SGD step uses lr=1.0 to keep FP32 master updates above rounding near unit-valued LayerNorm parameters. At lr=0.01, the root-layout update norm was about 4e-7, so a difference of one FP32 ULP exceeded the 5% relative gate. The forward/gradient comparisons and every acceptance threshold are unchanged.

The current modal-torch-latest workflow runs test selection on PR updates and executes its GPU suite on merge-queue entries. The historical full-CI timeout is not counted as a completed test run.

yh0903 and others added 7 commits August 31, 2026 10:53
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Accept explicitly disabled offload configs, observe the production eager boundary directly, and make small gradient and FP32 master-update parity checks discriminating.

Signed-off-by: yh0903 <helloyu0903@gmail.com>
@yh0903
yh0903 marked this pull request as ready for review September 5, 2026 22:10
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-05T22:15:31.846211Z 04cc82b Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 04cc82b71c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


# DeepSpeed Team

from __future__ import annotations

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add the required Signed-off-by trailer

This is a non-merge commit, but its message contains no Signed-off-by: trailer. The repository requires every non-merge commit to be created with --signoff, so this commit does not satisfy the merge requirements; recreate it using the configured author name and email with --signoff.

AGENTS.md reference: AGENTS.md:L8-L8

Useful? React with 👍 / 👎.

@tohtana tohtana left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for the improvement, @yh0903!
It's really good to have an official support to combine MoE and compile in DeepSpeed.

The change overall looks good to me. I left one comment. Here are two points I'd like to discuss:

(1) In terms of the interface to enable this feature, adding compile_mode to compile seems a bit too much as it is currently only for this feature. How about adding an option to compile section in the config like:

{
  "compile": {
    "autoep_non_moe": true
  }
}

I think we could also make this the default behavior.

(2) This feature currently compiles only decoder blocks, while excluding other layers including embedding and lm_heads. Can you explain the intension?

else:
module.forward = original_forward
for region, compiled_call in original_compiled_calls.items():
region._compiled_call_impl = compiled_call

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This also rejects a direct AutoEP child such as mlp, which is supported by custom AutoEP replacement. In that case, the separator is empty, but the parent is the callable model root (named_modules[""]). Only an empty module_name means the AutoEP layer itself is the root. Could we use if not module_name: here and add a regression test for a callable root with a direct AutoEP child?
Feel free to refer to yh0903#1 if it's useful.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the review! And sorry for the late response as I was working on the revision. I moved the opt-in into the DeepSpeed config as "compile": {"autoep_non_moe": true} and restored the existing engine.compile() signature. I kept the option disabled by default for this initial version so existing full-model compilation and configurations outside the supported AutoEP subset retain their current behavior.

I also fixed region discovery to allow a callable model root with a direct AutoEP child, following your suggestion. An AutoEP layer that is itself the model root remains rejected because it has no surrounding region. Regression coverage now exercises the root case through actual compilation and compares forward/backward results.

The decoder-region scope was intentional: repeated decoder blocks contain the attention, normalization, residual and dense work targeted by this optimization, while similar blocks can reuse compiled graphs. Embeddings and the lm_head are not inherently incompatible with compilation; they remain eager when outside the selected regions because extending the scope needs separate validation and measurement. A selected root region includes its non-MoE operations. I clarified this in the documentation.

The changes are in the latest commit with a follow-up test-fixture correction to keep single-step FP32 updates above rounding. All tolerances are unchanged. The final revision passes 57 targeted CPU tests and all eight Inductor parity cases on two H100s. Thanks!

Signed-off-by: yh0903 <helloyu0903@gmail.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
Signed-off-by: yh0903 <helloyu0903@gmail.com>
@yh0903 yh0903 changed the title Add an opt-in AutoEP non-MoE regional compile mode Add opt-in regional compilation around AutoEP layers Sep 21, 2026

@tohtana tohtana left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for the update, looks good to me.

@tohtana
tohtana enabled auto-merge September 22, 2026 06:23
@tohtana
tohtana added this pull request to the merge queue Sep 22, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 22, 2026
@tohtana
tohtana added this pull request to the merge queue Sep 22, 2026
Merged via the queue into deepspeedai:master with commit 6c2fae8 Sep 22, 2026
13 checks passed
@tohtana tohtana mentioned this pull request Sep 27, 2026
10 of 22 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants