Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
Mohsen Ghafoorian, Denis Korzhenkov, Amirhossein Habibian
摘要
Transformer-based video diffusion models (VDMs) deliver state-of-the-art video generation quality but are constrained by the quadratic cost of self-attention, making long sequences and high resolutions computationally expensive. While linear attention offers sub-quadratic complexity, previous approaches have failed to match the expressiveness of softmax attention unless retrained at significant computational cost. We introduce Attention Surgery, an efficient framework that enables linear or hybrid attention in pretrained VDMs, eliminating the need for training from scratch. Inspired by recent advances in language models, our method combines a novel hybrid attention mechanism-mixing softmax and linear tokens-with a lightweight distillation and fine-tuning pipeline requiring only a few GPU-days. Additionally, we incorporate a cost-aware block-rate strategy to balance expressiveness and efficiency across layers. Applied to Wan2.1 1.3B, a state-of-the-art efficient transformer VDM and evaluated on VBench, VBench2.0 and a human preference study, Attention Surgery achieves competitive results. Furthermore, measurements of on-mobile latency, memory usage, and FLOPs demonstrate notable improvements in scaling behavior for longer videos. Project page is available at: https://qualcomm-ai-research.github.io/attention-surgery.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion ModelsAritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amir Habibian 等ICLR 2026 · 被引用 15 次
- Neodragon: Mobile Video Generation Using Diffusion TransformerAnimesh Karnewar, Denis Korzhenkov, Ioannis Lelekas, Noor Fathima 等ICLR 2026 · 被引用 10 次
- ReHyAt: Recurrent Hybrid Attention for Video Diffusion TransformersMohsen Ghafoorian, Amirhossein HabibianCVPR 2026 · 被引用 5 次
- PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient InferenceDenis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian 等CVPR 2026 · 被引用 3 次
- VMonarch: Efficient Video Diffusion Transformers with Structured AttentionCheng Liang, Haoxian Chen, Liang Hou, Qi Fan 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper25
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen 等NeurIPS 2024 · 被引用 412 次
相关 Paper
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video GenerationYushi Huang, Xingtong Ge, Ruihao Gong, Chengtao Lv 等CVPR 2026 · 被引用 9 次
- Faster Video Diffusion with Trainable Sparse AttentionPeiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin 等NeurIPS 2025 · 被引用 6 次
- Veda: Scalable Video Diffusion via Distilled Sparse AttentionShihao Han, Hao Yang, Xiaofeng Mei, Xinting Hu 等ICML 2026 · 被引用 1 次
- EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward BackpropagationJiaqi Xu, Kunzhe Huang, Xinyi Zou, Yunkuo Chen 等ACM MM 2025
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao 等NeurIPS 2025 · 被引用 25 次
