Accelerating Diffusion Transformers with Token-wise Feature Caching
Chang Zou, Xuyang Liu, Ting Liu, Siteng Huang, Linfeng Zhang
摘要
Diffusion transformers have shown significant effectiveness in both image and video synthesis at the expense of huge computation costs. To address this problem, feature caching methods have been introduced to accelerate diffusion transformers by caching the features in previous timesteps and reusing them in the following timesteps. However, previous caching methods ignore that different tokens exhibit different sensitivities to feature caching, and feature caching on some tokens may lead to 10× more destruction to the overall generation quality compared with other tokens. In this paper, we introduce token-wise feature caching, allowing us to adaptively select the most suitable tokens for caching, and further enable us to apply different caching ratios to neural layers in different types and depths. Extensive experiments on PixArt-α, OpenSora, DiT and FLUX demonstrate our effectiveness in both image and video generation with no requirements for training. For instance, 2.36× and 1.93× acceleration are achieved on OpenSora and PixArt-α with almost no drop in generation quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper77
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei 等NeurIPS 2025 · 被引用 94 次
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang 等NeurIPS 2025 · 被引用 89 次
- DiCache: Let Diffusion Model Determine Its Own CacheJiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang 等ICLR 2026 · 被引用 36 次
- QVGen: Pushing the Limit of Quantized Video Generative ModelsYushi Huang, Ruihao Gong, Jing Liu, Yifu Ding 等ICLR 2026 · 被引用 22 次
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu 等ICLR 2026 · 被引用 20 次
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
相关 Paper
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li 等ICLR 2026 · 被引用 2 次
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature CachingZhixin Zheng, Xinyu Wang, Chang Zou, Shaobo Wang 等ACM MM 2025 · 被引用 4 次
- Diffusion on Demand: Selective Caching and Modulation for Efficient GenerationHee Min Choi, Hyoa Kang, Dokwan Oh, Nam Ik ChoNeurIPS 2025 · 被引用 2 次
- Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion TransformersShikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou 等AAAI 2026 · 被引用 10 次
- SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion TransformerTong Shao, Yusen Fu, Guoying Sun, Jingde Kong 等CVPR 2026 · 被引用 1 次
