Papers › Scaling Diffusion Transformers Efficiently via $μ$P

Scaling Diffusion Transformers Efficiently via $μ$P

21 May 2025arXiv:2505.15270archive 2025-07-28

Chenyu Zheng, Xinyu Zhang, Rongzhen Wang, Wei Huang, Zhi Tian, Weilin Huang, Jun Zhu, Chongxuan Li

Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization (μP) was proposed for vanilla Transformers, which enables stable HP transfer from small to large language models, and dramatically reduces tuning costs. However, it remains unclear whether μP of vanilla Transformers extends to diffusion Transformers, which differ architecturally and objectively. In this work, we generalize standard μP to diffusion Transformers and validate its effectiveness through large-scale experiments. First, we rigorously prove that μP of mainstream diffusion Transformers, including DiT, U-ViT, PixArt-α, and MMDiT, aligns with that of the vanilla Transformer, enabling the direct application of existing μP methodologies. Leveraging this result, we systematically demonstrate that DiT-μP enjoys robust HP transferability. Notably, DiT-XL-2-μP with transferred learning rate achieves 2.9 times faster convergence than the original DiT-XL-2. Finally, we validate the effectiveness of μP on text-to-image generation by scaling PixArt-α from 0.04B to 0.61B and MMDiT from 0.18B to 18B. In both cases, models under μP outperform their respective baselines while requiring small tuning cost, only 5.5% of one training run for PixArt-α and 3% of consumption by human experts for MMDiT-18B. These results establish μP as a principled and efficient framework for scaling diffusion Transformers.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ML-GSAI/Scaling-Diffusion-Transformers-muP officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDiffusionDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections