Attention-Driven Training-Free Efficiency Enhancement of Diffusion Models
Hongjie Wang, Difan Liu, Yan Kang, Yijun Li, Zhe Lin, Niraj K. Jha, Yuchen Liu
Abstract
Diffusion Models (DMs) have exhibited superior performance in generating high-quality and diverse images. However, this exceptional performance comes at the cost of expensive architectural design, particularly due to the attention module heavily used in leading models. Existing works mainly adopt a retraining process to enhance DM efficiency. This is computationally expensive and not very scalable. To this end, we introduce the Attention-driven Training-free Efficient Diffusion Model (AT-EDM) framework that leverages attention maps to perform run-time pruning of redundant tokens, without the need for any retraining. Specifically, for single-denoising-step pruning, we develop a novel ranking algorithm, Generalized Weighted Page Rank (G-WPR), to identify redundant tokens, and a similarity-based recovery method to restore tokens for the convolution operation. In addition, we propose a Denoising-Steps-Aware Pruning (DSAP) approach to adjust the pruning budget across different denoising timesteps for better generation quality. Extensive evaluations show that AT-EDM performs favorably against prior art in terms of efficiency (e.g., 38.8% FLOPs saving and up to 1.53× speed-up over Stable Diffusion XL) while maintaining nearly the same FID and CLIP scores as the full model. Project webpage: https://atedm.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8615c3a-e35b-4927-a60c-925970e0caacCited by top-tier papers23
- MagCache: Fast Video Generation with Magnitude-Aware CacheZehong Ma, Longhui Wei, Feng Wang, Shiliang Zhang et al.NeurIPS 2025 · 41 citations
- Training-Free and Hardware-Friendly Acceleration for Diffusion Models via Similarity-based Token PruningEvelyn Zhang, Jiayi Tang, Xuefei Ning, Linfeng ZhangAAAI 2025 · 36 citations
- PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video GenerationJiangshan Wang, Kang Zhao, Jiayi Guo, Jiayu Wang et al.ICLR 2026 · 6 citations
- FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningHang Guo, Yawei Li, Taolin Zhang, Jiangshan Wang et al.ICCV 2025 · 5 citations
- RAPID: Tri-Level Reinforced Acceleration Policies for Diffusion TransformerWangbo Zhao, Yizeng Han, Zhiwei Tang, Jiasheng Tang et al.ICLR 2026 · 5 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationHaipeng Fang, Sheng Tang, Juan Cao, Enshuo Zhang et al.CVPR 2025
- Structural Pruning for Diffusion ModelsGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2023 · 257 citations
- F³-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video SynthesisSitong Su, Jianzhi Liu, Lianli Gao, Jingkuan SongAAAI 2024
- EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like SketchingXinwang Chen, Ning Liu, Yichen Zhu, Feifei Feng et al.NeurIPS 2024 · 5 citations
- Learnable Sparsity for Vision Generative ModelsYang Zhang, Er Jin, Wenzhong Liang, Yanfei Dong et al.ICLR 2026 · 8 citations
