FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
Jingfeng Yao, Cheng Wang, Wenyu Liu, Xinggang Wang
Abstract
Diffusion Transformers (DiT) have attracted significant attention in research. However, they suffer from a slow convergence rate. In this paper, we aim to accelerate DiT training without any architectural modification. We identify the following issues in the training process: firstly, certain training strategies do not consistently perform well across different data. Secondly, the effectiveness of supervision at specific timesteps is limited. In response, we propose the following contributions:
(1) We introduce a new perspective for interpreting the failure of the strategies. Specifically, we slightly extend the definition of Signal-to-Noise Ratio (SNR) and suggest observing the Probability Density Function (PDF) of SNR to understand the essence of the data robustness of the strategy. (2) We conduct numerous experiments and report over one hundred experimental results to empirically summarize a unified accelerating strategy from the perspective of PDF. (3) We develop a new supervision method that further accelerates the training process of DiT. Based on them, we propose FasterDiT, an exceedingly simple and practicable design strategy. With few lines of code modifications, it achieves 2.30 FID on ImageNet at 256×256 resolution with 1000 iterations, which is comparable to DiT (2.27 FID) but 7× faster in training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb2f664c-a3ef-4a76-b82f-be3c62b2ae9bCited by top-tier papers34
- Pixel-Perfect Depth with Semantics-Prompted Diffusion TransformersGangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang et al.NeurIPS 2025 · 58 citations
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris et al.NeurIPS 2025 · 47 citations
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingZiqiao Wang, Wangbo Zhao, Yuhao Zhou, Zekai Li et al.NeurIPS 2025 · 37 citations
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive GenerationAnlin Zheng, Xin Wen, Xuanyang Zhang, Chuofan Ma et al.NeurIPS 2025 · 21 citations
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- SD-DiT: Unleashing the Power of Self-Supervised Discrimination in Diffusion Transformer*Rui Zhu, Yingwei Pan, Yehao Li, Ting Yao et al.CVPR 2024 · 15 citations
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 102 citations
- Efficient Diffusion Training via Min-SNR Weighting StrategyTiankai Hang, Shuyang Gu, Chen Li, Jianmin Bao et al.ICCV 2023 · 261 citations
- MC-DiT: Contextual Enhancement via Clean-to-Clean Reconstruction for Masked Diffusion ModelsGuanghao Zheng, Yuchen Liu, Wenrui Dai, Chenglin Li et al.NeurIPS 2024 · 2 citations
- LiT: Delving into a Simple Linear Diffusion Transformer for Image GenerationJiahao Wang, Ning Kang, Lewei Yao, Mengzhao Chen et al.ICCV 2025 · 10 citations
