Training-Free and Hardware-Friendly Acceleration for Diffusion Models via Similarity-based Token Pruning
Evelyn Zhang, Jiayi Tang, Xuefei Ning, Linfeng Zhang
Abstract
The excellent performance of diffusion models in image generation is always accompanied by overlarge computation costs, which have prevented the application of diffusion models in edge devices and interactive applications. Previous works mainly focus on using fewer sampling steps and compressing the denoising network of diffusion models, while this paper proposes to accelerate diffusion models by introducing SiTo, a similarity-based token pruning method that adaptive prunes the redundant tokens in the input data. SiTo is designed to maximize the similarity between model prediction with and without token pruning by using cheap and hardware-friendly operations, leading to significant acceleration ratios without performance drop, and even sometimes improvements in the generation quality. For instance, the zero-shot evaluation shows SiTo leads to 1.90x and 1.75x acceleration on COCO30K and ImageNet with 1.33 and 1.15 FID reduction at the same time. Besides, SiTo has no training requirements and does not require any calibration data, making it plug-and-play in real-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42a79ef5-3f36-4ab1-b492-a201be39b03fCited by top-tier papers17
- From Reusing to Forecasting: Accelerating Diffusion Models With TaylorseersJiacheng Liu, Chang Zou, Yuanhuiyi Lyu, Junjie Chen et al.ICCV 2025 · 12 citations
- Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion TransformersShikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou et al.AAAI 2026 · 10 citations
- SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion ModelsJiwoo Chung, Sangeek Hyun, MinKyu Lee, Byeongju Han et al.CVPR 2026 · 9 citations
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion TransformersShikang Zheng, Guantao Chen, Qinming Zhou, Yuqi Lin et al.ICLR 2026 · 7 citations
- PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video GenerationJiangshan Wang, Kang Zhao, Jiayi Guo, Jiayu Wang et al.ICLR 2026 · 6 citations
Builds on16
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
Related papers
- DiP-GO: A Diffusion Pruner via Few-step Gradient OptimizationHaowei Zhu, Dehua Tang, Ji Liu, Mingjie Lu et al.NeurIPS 2024 · 51 citations
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationHaipeng Fang, Sheng Tang, Juan Cao, Enshuo Zhang et al.CVPR 2025
- SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion TransformerTong Shao, Yusen Fu, Guoying Sun, Jingde Kong et al.CVPR 2026 · 1 citation
- Attention-Driven Training-Free Efficiency Enhancement of Diffusion ModelsHongjie Wang, Difan Liu, Yan Kang, Yijun Li et al.CVPR 2024 · 4 citations
- Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained TransformersHongjie Wang, Bhishma Dedhia, Niraj K. JhaCVPR 2024 · 24 citations
