ResCa: Residual Caching for Diffusion Transformers Acceleration
Haipeng Fang, Yu Li, Fan Tang, Yixing Lu, Juan Cao, Sheng Tang
Abstract
Diffusion transformers have achieved remarkable progress in high-quality image and video generation, but their high computational overhead remains a significant issue. Existing token reduction-based acceleration techniques, such as caching and merging, attempt to reduce this cost from both temporal and spatial perspectives but often compromise generation quality by introducing non-updated or nonself denoising directions. In this paper, we propose Residual Caching (ResCa), a novel, training-free framework that introduces a proxy denoising perspective to overcome these limitations. ResCa achieves acceleration while maintaining a denoising trajectory that is both self and updated. The core idea is to perform true denoising on only one "proxy token" within each trajectory-based cluster, and use its computed multi-order residuals to guide the "simulated denoising" of all other tokens. ResCa can be seamlessly integrated into various diffusion models, including DiT, FLUX, and HunyuanVideo. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method, achieving up to a 5.5× acceleration in GFLOPs while maintaining near-lossless generation quality on FLUX. Project page: https://fanghaipeng.github.io/ResCa.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f97511b5-c8b7-4974-bfd8-15230ec66f05Builds on33
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
Related papers
- ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval CachingLihui Gu, Jingbin He, Lianghao Su, Kang He et al.ICLR 2026
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature CachingZhixin Zheng, Xinyu Wang, Chang Zou, Shaobo Wang et al.ACM MM 2025 · 4 citations
- Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheBowen Cui, Yuanbin Wang, Huajiang Xu, Biaolong Chen et al.CVPR 2026 · 6 citations
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion TransformersShikang Zheng, Guantao Chen, Qinming Zhou, Yuqi Lin et al.ICLR 2026 · 7 citations
- Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Diffusion TransformersGuantao Chen, Shikang Zheng, Yuqi Lin, Linfeng ZhangCVPR 2026
