Scaling Diffusion Transformers Efficiently via μP
Chenyu Zheng, Xinyu Zhang, Rongzhen Wang, Wei Huang, Zhi Tian, Weilin Huang, Jun Zhu, Chongxuan Li
Abstract
Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization (P) was proposed for vanilla Transformers, which enables stable HP transfer from small to large language models, and dramatically reduces tuning costs. However, it remains unclear whether P of vanilla Transformers extends to diffusion Transformers, which differ architecturally and objectively. In this work, we generalize standard P to diffusion Transformers and validate its effectiveness through large-scale experiments. First, we rigorously prove that P of mainstream diffusion Transformers, including U-ViT, DiT, PixArt-, and MMDiT, aligns with that of the vanilla Transformer, enabling the direct application of existing P methodologies. Leveraging this result, we systematically demonstrate that DiT-P enjoys robust HP transferability. Notably, DiT-XL-2-P with transferred learning rate achieves 2.9 times faster convergence than the original DiT-XL-2. Finally, we validate the effectiveness of P on text-to-image generation by scaling PixArt- from 0.04B to 0.61B and MMDiT from 0.18B to 18B. In both cases, models under P outperform their respective baselines while requiring small tuning cost, only 5.5% of one training run for PixArt- and 3% of consumption by human experts for MMDiT-18B. These results establish P as a principled and efficient framework for scaling diffusion Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01a48dd7-26fd-404c-aa5f-a2f17ae5b7baCited by top-tier papers1
Ask how each one uses itBuilds on35
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick et al.ICCV 2025 · 9 citations
- EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingHaotian Sun, Tao Lei, Bowen Zhang, Yanghao Li et al.ICLR 2025
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi et al.ICLR 2026 · 11 citations
- LiT: Delving into a Simple Linear Diffusion Transformer for Image GenerationJiahao Wang, Ning Kang, Lewei Yao, Mengzhao Chen et al.ICCV 2025 · 10 citations
