Skip to content

MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators #14882

Description

@li-lizhe

Describe the bug

MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:

with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
    video = components.vae.decode(latents, return_dict=False)[0]

The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.

Suggested change - gate on "not CPU" instead of "CUDA only":

enabled=device.type != "cpu",

CPU keeps the current behaviour (autocast disabled).

Reproduction

On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:

import torch

x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
    print((x @ x).dtype)   # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
    print((x @ x).dtype)   # torch.float32

Not measured

End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.

Related but distinct

#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.

System Info

  • diffusers: main
  • hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu); applies to any non-CUDA accelerator.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions