Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
Shikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou, Peiliang Cai, Chang Zou, Jiacheng Liu, Yuqi Lin, Junjie Chen, Yue Ma, Linfeng Zhang
Abstract
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature caching techniques have been proposed to accelerate inference by reusing hidden representations from previous timesteps. However, current methods often struggle to maintain generation quality at high acceleration ratios, where prediction errors increase sharply due to the inherent instability of long-step forecasting. In this work, we adopt an ordinary differential equation (ODE) perspective on the hiddenfeature sequence, modeling layer representations along the trajectory as a feature-ODE. We attribute the degradation of existing caching strategies to their inability to robustly integrate historical features under large skipping intervals. To address this, we propose FoCa (Forecast-then-Calibrate), which treats feature caching as a feature-ODE solving problem. Extensive experiments on image synthesis, video generation, and super-resolution tasks demonstrate the effectiveness of FoCa, especially under aggressive acceleration. Without additional training, FoCa achieves near-lossless speedups of 5.50× on FLUX, 6.45× on HunyuanVideo, 3.17× on Inf-DiT, and maintains high quality with a 4.53× speedup on DiT. Our code will be released upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 335d9228-2548-4cf1-b8a4-6dd9d4434d06Cited by top-tier papers8
- From Sketch to Fresco: Efficient Diffusion Transformer with Progressive ResolutionShikang Zheng, Guantao Chen, Landis He, Jiacheng Liu et al.CVPR 2026 · 6 citations
- LESA: Learnable Stage-Aware Predictors for Diffusion Model AccelerationPeiliang Cai, Jiacheng Liu, Haowen Xu, Xinyu Wang et al.CVPR 2026 · 4 citations
- TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion AccelerationHaowei Zhu, Tingxuan Huang, Xing Wang, Tianyu Zhao et al.CVPR 2026 · 3 citations
- FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent PredictionShuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han et al.CVPR 2026 · 3 citations
- Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion ModelsZhirong Shen, Rui Huang, Jiacheng Liu, Chang Zou et al.CVPR 2026 · 1 citation
Builds on16
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Structural Pruning for Diffusion ModelsGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2023 · 257 citations
Related papers
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion TransformersShikang Zheng, Guantao Chen, Qinming Zhou, Yuqi Lin et al.ICLR 2026 · 7 citations
- From Reusing to Forecasting: Accelerating Diffusion Models With TaylorseersJiacheng Liu, Chang Zou, Yuanhuiyi Lyu, Junjie Chen et al.ICCV 2025 · 12 citations
- Accelerating Diffusion Transformers with Token-wise Feature CachingChang Zou, Xuyang Liu, Ting Liu, Siteng Huang et al.ICLR 2025
- ResCa: Residual Caching for Diffusion Transformers AccelerationHaipeng Fang, Yu Li, Fan Tang, Yixing Lu et al.CVPR 2026
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature CachingZhixin Zheng, Xinyu Wang, Chang Zou, Shaobo Wang et al.ACM MM 2025 · 4 citations
