Tread: Token Routing for Efficient Architecture-Agnostic Diffusion Training
Felix Krause, Timy Phan, Ming Gui, Stefan Andreas Baumann, Vincent Tao Hu, Björn Ommer
Abstract
Diffusion models have emerged as the mainstream approach for visual generation. However, these models typically suffer from sample inefficiency and high training costs. Consequently, methods for efficient finetuning, inference and personalization were quickly adopted by the community. However, training these models in the first place remains very costly. While several recent approaches - including masking, distillation, and architectural modifications - have been proposed to improve training efficiency, each of these methods comes with a tradeoff: they achieve enhanced performance at the expense of increased computational cost or vice versa. In contrast, this work aims to improve training efficiency as well as generative performance at the same time through routes that act as a transport mechanism for randomly selected tokens from early layers to deeper layers of the model. Our method is not limited to the common transformer-based model - it can also be applied to state-space models and achieves this without architectural modifications or additional parameters. Finally, we show that TREAD reduces computational cost and simultaneously boosts model performance on the standard ImageNet-256 benchmark in class-conditional synthesis. Both of these benefits multiply to a convergence speedup of 14x at 400K training iterations compared to DiT and 37x compared to the best benchmark performance of DiT at 7M training iterations. Furthermore, we achieve a competitive FID of 2.09 in a guided and 3.93 in an unguided setting, which improves upon the DiT, without architectural changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45161419-8384-47d0-bb28-da32dd565c36Cited by top-tier papers8
- Representation Alignment for Diffusion Transformers without External ComponentsDengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang et al.ICLR 2026 · 532 citations
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingZiqiao Wang, Wangbo Zhao, Yuhao Zhou, Zekai Li et al.NeurIPS 2025 · 37 citations
- Guiding a Diffusion Transformer with the Internal Dynamics of ItselfXingyu Zhou, Qifan Li, Xiaobin Hu, Hai Chen et al.CVPR 2026 · 13 citations
- SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion TransformersDogyun Park, Moayed Haji-Ali, Yanyu Li, Willi Menapace et al.ICLR 2026 · 6 citations
- One Model, Many Budgets: Elastic Latent Interfaces for Diffusion TransformersMoayed Haji Ali, Willi Menapace, Ivan Skorokhodov, Dogyun Park et al.CVPR 2026 · 4 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Language-Guided Image Tokenization for GenerationKaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross et al.CVPR 2025
- Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion ModelsJingfeng Yao, Bin Yang, Xinggang WangCVPR 2025
- EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like SketchingXinwang Chen, Ning Liu, Yichen Zhu, Feifei Feng et al.NeurIPS 2024 · 5 citations
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion ModelingYuang Ai, Qihang Fan, Xuefeng Hu, Zhenheng Yang et al.NeurIPS 2025 · 8 citations
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li et al.ICLR 2026 · 2 citations
