DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
Haowei Zhu, Ji Liu, Ziqiong Liu, Dong Li, Jun-Hai Yong, Bin Wang, Emad Barsoum
Abstract
Diffusion models demonstrate outstanding performance in image generation, but their multi-step inference mechanism requires immense computational cost. Previous works accelerate inference by leveraging layer or token cache techniques to reduce computational cost. However, these methods fail to achieve superior acceleration performance in few-step diffusion transformer models due to inefficient feature caching strategies, manually designed sparsity allocation, and the practice of retaining complete forward computations in several steps in these token cache methods. To tackle these challenges, we propose a differentiable layer-wise sparsity optimization framework for diffusion transformer models, leveraging token caching to reduce token computation costs and enhance acceleration. Our method optimizes layer-wise sparsity allocation in an end-to-end manner through a learnable network combined with a dynamic programming solver. Additionally, our proposed two-stage training strategy eliminates the need for full-step processing in existing methods, further improving efficiency. We conducted extensive experiments on a range of diffusion-transformer models, including DiT-XL/2, PixArt-, FLUX, and Wan2.1. Across these architectures, our method consistently improves efficiency without degrading sample quality. For example, on PixArt- with 20 sampling steps, we reduce computational cost by while achieving generation metrics that surpass those of the original model, substantially outperforming prior approaches. These results demonstrate that our method delivers large efficiency gains while often improving generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9928dade-9a0a-41d5-8ecf-56af3f709910Builds on36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Accelerating Diffusion Transformers with Token-wise Feature CachingChang Zou, Xuyang Liu, Ting Liu, Siteng Huang et al.ICLR 2025
- SparseDiT: Token Sparsification for Efficient Diffusion TransformerShuning Chang, Pichao Wang, Jiasheng Tang, Fan Wang et al.NeurIPS 2025 · 9 citations
- Diffusion on Demand: Selective Caching and Modulation for Efficient GenerationHee Min Choi, Hyoa Kang, Dokwan Oh, Nam Ik ChoNeurIPS 2025 · 2 citations
- Elastic Diffusion TransformerJiangshan Wang, Zeqiang Lai, Jiarui Chen, Jiayi Guo et al.ICML 2026 · 7 citations
- ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer AccelerationFanpu Cao, Yaofo Chen, Zeng You, Wei LuoAAAI 2026
