Accelerating Diffusion Transformer via Gradient-Optimized Cache
Junxiang Qiu, Lin Liu, Shuo Wang, Jinda Lu, Kezhou Chen, Yanbin Hao
Abstract
Feature caching has emerged as an effective strategy to accelerate diffusion transformer (DiT) sampling through temporal feature reuse. It is a challenging problem since (1) Progressive error accumulation from cached blocks significantly degrades generation quality, particularly when over 50% of blocks are cached; (2) Current error compensation approaches neglect dynamic perturbation patterns during the caching process, leading to suboptimal error correction. To solve these problems, we propose the Gradient-Optimized Cache (GOC) with two key innovations: (1) Cached Gradient Propagation: A gradient queue dynamically computes the gradient differences between cached and recomputed features. These gradients are weighted and propagated to subsequent steps, directly compensating for the approximation errors introduced by caching. (2) Inflection-Aware Optimization: Through statistical analysis of feature variation patterns, we identify critical inflection points where the denoising trajectory changes direction. By aligning gradient updates with these detected phases, we prevent conflicting gradient directions during error correction. Extensive evaluations on ImageNet demonstrate GOC's superior trade-off between efficiency and quality. With 50% cached blocks, GOC achieves IS 216.28 () and FID 3.907 () compared to baseline DiT, while maintaining identical computational costs. These improvements persist across various cache ratios, demonstrating robust adaptability to different acceleration requirements. Code is available at https://github.com/qiujx0520/GOC_ICCV2025.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Enhancing CLIP Robustness via Cross-Modality AlignmentXingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao et al.NeurIPS 2025 · 17 citations
- Astraea: A Token-wise Acceleration Framework for Video Diffusion TransformersHaosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu et al.ICLR 2026 · 9 citations
- Relational Feature Caching for Accelerating Diffusion TransformersByunggwan Son, Jeimin Jeon, Jeongwoo Choi, Bumsub HamICLR 2026 · 6 citations
- Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheBowen Cui, Yuanbin Wang, Huajiang Xu, Biaolong Chen et al.CVPR 2026 · 6 citations
- Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward ModelYuan Wang, Borui Liao, Huijuan Huang, Jinda Lu et al.CVPR 2026 · 5 citations
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
Related papers
- Accelerating Diffusion Transformer via Error-Optimized CacheJunxiang Qiu, Shuo Wang, Jinda Lu, Lin Liu et al.ACM MM 2025 · 3 citations
- Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion TransformersShikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou et al.AAAI 2026 · 10 citations
- ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer AccelerationFanpu Cao, Yaofo Chen, Zeng You, Wei LuoAAAI 2026
- Revisiting Redundancy in Diffusion Transformers: A Temporal-Spatial Joint Caching Strategy for Efficient SamplingChenxi Du, Yongheng Deng, Ju Ren, Yaoxue ZhangKDD 2026
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error MinimizationTong Shao, Yusen Fu, Guoying Sun, Jingde Kong et al.ICLR 2026 · 2 citations
