Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
Shikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou, Peiliang Cai, Chang Zou, Jiacheng Liu, Yuqi Lin, Junjie Chen, Yue Ma, Linfeng Zhang
摘要
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature caching techniques have been proposed to accelerate inference by reusing hidden representations from previous timesteps. However, current methods often struggle to maintain generation quality at high acceleration ratios, where prediction errors increase sharply due to the inherent instability of long-step forecasting. In this work, we adopt an ordinary differential equation (ODE) perspective on the hiddenfeature sequence, modeling layer representations along the trajectory as a feature-ODE. We attribute the degradation of existing caching strategies to their inability to robustly integrate historical features under large skipping intervals. To address this, we propose FoCa (Forecast-then-Calibrate), which treats feature caching as a feature-ODE solving problem. Extensive experiments on image synthesis, video generation, and super-resolution tasks demonstrate the effectiveness of FoCa, especially under aggressive acceleration. Without additional training, FoCa achieves near-lossless speedups of 5.50× on FLUX, 6.45× on HunyuanVideo, 3.17× on Inf-DiT, and maintains high quality with a 4.53× speedup on DiT. Our code will be released upon acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- From Sketch to Fresco: Efficient Diffusion Transformer with Progressive ResolutionShikang Zheng, Guantao Chen, Landis He, Jiacheng Liu 等CVPR 2026 · 被引用 6 次
- LESA: Learnable Stage-Aware Predictors for Diffusion Model AccelerationPeiliang Cai, Jiacheng Liu, Haowen Xu, Xinyu Wang 等CVPR 2026 · 被引用 4 次
- TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion AccelerationHaowei Zhu, Tingxuan Huang, Xing Wang, Tianyu Zhao 等CVPR 2026 · 被引用 3 次
- FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent PredictionShuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han 等CVPR 2026 · 被引用 3 次
- Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion ModelsZhirong Shen, Rui Huang, Jiacheng Liu, Chang Zou 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper16
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Structural Pruning for Diffusion ModelsGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2023 · 被引用 257 次
相关 Paper
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion TransformersShikang Zheng, Guantao Chen, Qinming Zhou, Yuqi Lin 等ICLR 2026 · 被引用 7 次
- From Reusing to Forecasting: Accelerating Diffusion Models With TaylorseersJiacheng Liu, Chang Zou, Yuanhuiyi Lyu, Junjie Chen 等ICCV 2025 · 被引用 12 次
- Accelerating Diffusion Transformers with Token-wise Feature CachingChang Zou, Xuyang Liu, Ting Liu, Siteng Huang 等ICLR 2025
- ResCa: Residual Caching for Diffusion Transformers AccelerationHaipeng Fang, Yu Li, Fan Tang, Yixing Lu 等CVPR 2026
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature CachingZhixin Zheng, Xinyu Wang, Chang Zou, Shaobo Wang 等ACM MM 2025 · 被引用 4 次
