APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
Yuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu
Abstract
The efficiency of long-video inference remains a critical bottleneck, mainly due to the dense computation in the prefill stage of Large Multimodal Models (LMMs). Existing methods either compress visual embeddings or apply sparse attention on a single GPU, yielding limited acceleration or degraded performance and restricting LMMs from handling longer, more complex videos. To overcome these issues, we propose APB-V, a sequence-parallel framework with optimized attention that accelerates long-video inference across multiple GPUs. By distributing approximate attention, APB-V reduces computation and increases parallelism, enabling efficient processing of more visual embeddings without compression and thereby improving task performance. System-level optimizations, such as load balancing and fused forward passes, further unleash the potential of APB-V, delivering speedups of 12.72×, 1.70×, and 1.18× over FLASHATTN, ZIGZAGRING, and APB, without notable performance loss. Code available at https://github.com/thunlp/APB .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 748db597-b5ac-4f20-a52c-11d348e697b8Builds on18
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheZirui Liu, Jiayi Yuan, Hongye Jin, Shaochen (Henry) Zhong et al.ICML 2024 · 436 citations
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
- Don't Look Twice: Faster Video Transformers with Run-Length TokenizationRohan Choudhury, Guanglei Zhu, Sihan Liu, Koichiro Niinuma et al.NeurIPS 2024 · 56 citations
- Video Super-Resolution Transformer with Masked Inter&Intra-Frame AttentionXingyu Zhou, Leheng Zhang, Xiaorui Zhao, Keze Wang et al.CVPR 2024 · 20 citations
Related papers
- APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUsYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao et al.ACL 2025
- LongVILA: Scaling Long-Context Visual Language Models for Long VideosYukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu et al.ICLR 2025 · 1 citation
- LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language ModelsTzu-Tao Chang, Shivaram VenkataramanICML 2025
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM InferenceSamir Khaki, Junxian Guo, Jiaming Tang, Shang Yang et al.ICCV 2025 · 3 citations
- Free-Moref: Instantly Multiplexing Context Perception Capabilities of Video-Mllms Within Single InferenceKuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang et al.ICCV 2025
