ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval Caching
Lihui Gu, Jingbin He, Lianghao Su, Kang He, Wenxiao Wang, Yuliang Liu
Abstract
Diffusion Transformers (DiTs) have emerged as powerful generative models, but their iterative denoising structure and deep transformer blocks incur substantial computational overhead, limiting the accessibility and practical deployment of highquality video generation. To address this bottleneck, we propose ScalingCache, a training-free acceleration framework specifically designed for DiTs. Scaling-Cache exploits the inherent redundancy in model representations by performing lightweight offline analysis on a small number of samples and dynamically reusing previously computed activations during inference, thereby avoiding full computation at certain denoising steps. Experimental results demonstrate that ScalingCache achieves significant acceleration in both image and video generation tasks while maintaining near-lossless generation quality. On widely used video generation models including Wan2.1 and HunyuanVideo, it achieves approximately 2.5× acceleration with only 0.5% drop in VBench scores; on FLUX, it achieves 3.1× near-lossless acceleration, with human preference tests showing comparable quality to original outputs. Moreover, under similar acceleration ratios, ScalingCache outperforms prior state-of-the-art caching strategies, achieving a 45% reduction in LPIPS for text-to-image generation and 20-30% reduction for text-to-video generation, highlighting its superior fidelity preservation. Our code is available at https://github.com/KlingAIResearch/ScalingCache .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bb3da74-5d7e-4484-ac94-6cf468b000fcBuilds on17
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Learning-to-Cache: Accelerating Diffusion Transformer via Layer CachingXinyin Ma, Gongfan Fang, Michael Bi Mi, Xinchao WangNeurIPS 2024 · 167 citations
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li et al.NeurIPS 2025 · 114 citations
Related papers
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise CachingHanshuai Cui, Zhiqing Tang, Zhifei Xu, Zhi Yao et al.ICLR 2026 · 11 citations
- Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheBowen Cui, Yuanbin Wang, Huajiang Xu, Biaolong Chen et al.CVPR 2026 · 6 citations
- Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising TimestepTianyi Liu, Ye Lu, Linfeng Zhang, Chen Cai et al.CVPR 2026 · 2 citations
- ResCa: Residual Caching for Diffusion Transformers AccelerationHaipeng Fang, Yu Li, Fan Tang, Yixing Lu et al.CVPR 2026
- Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Diffusion TransformersGuantao Chen, Shikang Zheng, Yuqi Lin, Linfeng ZhangCVPR 2026
