Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value Decomposition
Zhiyuan Chen, Keyi Li, Yifan Jia, Le Ye, Yufei Ma
Abstract
Diffusion transformer (DiT) models have achieved remarkable success in image generation, thanks for their exceptional generative capabilities and scalability. Nonetheless, the iterative nature of diffusion models (DMs) results in high computation complexity, posing challenges for deployment. Although existing cache-based acceleration methods try to utilize the inherent temporal similarity to skip redundant computations of DiT, the lack of correction may induce potential quality degradation. In this paper, we propose increment-calibrated caching, a training-free method for DiT acceleration, where the calibration parameters are generated from the pre-trained model itself with low-rank approximation. To deal with the possible correction failure arising from outlier activations, we introduce channelaware Singular Value Decomposition (SVD), which further strengthens the calibration effect. Experimental results show that our method always achieve better performance than existing naive caching methods with a similar computation resource budget. When compared with 35-step DDIM, our method eliminates more than 45% computation and improves IS by 12 at the cost of less than 0.06 FID increase.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- ERTACache: Error Rectification and Timesteps Adjustment for Efficient DiffusionXurui Peng, Chenqian Yan, Hong Liu, Rui Ma et al.ICLR 2026 · 10 citations
- Fast3Dcache: Training-free 3D Geometry Synthesis AccelerationMengyu Yang, Yanming Yang, Chenyi Xu, Chenxi Song et al.CVPR 2026 · 4 citations
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error MinimizationTong Shao, Yusen Fu, Guoying Sun, Jingde Kong et al.ICLR 2026 · 2 citations
- RSTR: Reducing SpatioTemporal Redundancy in Diffusion TransformersRuitong Sun, Tianze Yang, Wei Niu, Jin SunICML 2026 · 2 citations
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Diffusion TransformersGuantao Chen, Shikang Zheng, Yuqi Lin, Linfeng ZhangCVPR 2026
- ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval CachingLihui Gu, Jingbin He, Lianghao Su, Kang He et al.ICLR 2026
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise CachingHanshuai Cui, Zhiqing Tang, Zhifei Xu, Zhi Yao et al.ICLR 2026 · 11 citations
- Adaptive Caching for Faster Video Generation With Diffusion TransformersKumara Kahatapitiya, Haozhe Liu, Sen He, Ding Liu et al.ICCV 2025 · 5 citations
- Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheBowen Cui, Yuanbin Wang, Huajiang Xu, Biaolong Chen et al.CVPR 2026 · 6 citations
