Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models
Zhirong Shen, Rui Huang, Jiacheng Liu, Chang Zou, Peiliang Cai, Shikang Zheng, Zhengyi Shi, Liang Feng, Linfeng Zhang
Abstract
Diffusion Transformers (DiTs) have achieved state-of-theart image and video generation performance, but sampling remains expensive due to repeated transformer forward passes over many timesteps. Feature caching offers a training-free way to accelerate inference by reusing or forecasting hidden representations, yet recent forecastingbased methods derive their coefficients from hand-crafted formulas (e.g., Taylor expansion), which ultimately reduce to fixed linear combinations of a few historical features. Such fixed coefficients are suboptimal and fragile under aggressive skipping. In this paper, we first show that existing forecasting-based caching methods can be unified in a common linear form, and then analyze DiT feature trajectories, finding that for most denoising steps the current feature can be reconstructed from past features with projection fidelity above 0.95, indicating that accurate linear prediction is feasible. Motivated by this, we propose L 2 P (Learnable Linear Predictor), a simple data-driven caching framework that replaces hand-designed coefficients with learnable per-timestep weights trained on a small set of cached trajectories using a mean-squared error loss, converging in about 20 seconds on a single GPU. Extensive experiments on state-of-the-art DiTs demonstrate that L 2 P consistently outperforms existing caching baselines: on FLUX.1-dev, L 2 P achieves a 4.55× FLOPs reduction and 4.15× latency speedup with a PSNR of 31.459, and on Qwen-Image and Qwen-Image-Lightning, it maintains high visual fidelity even under up to 7.18× acceleration, where prior methods suffer from noticeable quality degradation. These results show that learning linear predictors is a practical and effective alternative to designing increasingly complex forecasting formulas for efficient diffusion
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
Related papers
- LESA: Learnable Stage-Aware Predictors for Diffusion Model AccelerationPeiliang Cai, Jiacheng Liu, Haowen Xu, Xinyu Wang et al.CVPR 2026 · 4 citations
- Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion TransformersShikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou et al.AAAI 2026 · 10 citations
- Adaptive Spectral Feature Forecasting for Diffusion Sampling AccelerationJiaqi Han, Juntong Shi, Puheng Li, Haotian Ye et al.CVPR 2026 · 4 citations
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise CachingHanshuai Cui, Zhiqing Tang, Zhifei Xu, Zhi Yao et al.ICLR 2026 · 11 citations
- Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheBowen Cui, Yuanbin Wang, Huajiang Xu, Biaolong Chen et al.CVPR 2026 · 6 citations
