APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
Yuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu
摘要
The efficiency of long-video inference remains a critical bottleneck, mainly due to the dense computation in the prefill stage of Large Multimodal Models (LMMs). Existing methods either compress visual embeddings or apply sparse attention on a single GPU, yielding limited acceleration or degraded performance and restricting LMMs from handling longer, more complex videos. To overcome these issues, we propose APB-V, a sequence-parallel framework with optimized attention that accelerates long-video inference across multiple GPUs. By distributing approximate attention, APB-V reduces computation and increases parallelism, enabling efficient processing of more visual embeddings without compression and thereby improving task performance. System-level optimizations, such as load balancing and fused forward passes, further unleash the potential of APB-V, delivering speedups of 12.72×, 1.70×, and 1.18× over FLASHATTN, ZIGZAGRING, and APB, without notable performance loss. Code available at https://github.com/thunlp/APB .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheZirui Liu, Jiayi Yuan, Hongye Jin, Shaochen (Henry) Zhong 等ICML 2024 · 被引用 436 次
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 被引用 205 次
- Don't Look Twice: Faster Video Transformers with Run-Length TokenizationRohan Choudhury, Guanglei Zhu, Sihan Liu, Koichiro Niinuma 等NeurIPS 2024 · 被引用 56 次
- Video Super-Resolution Transformer with Masked Inter&Intra-Frame AttentionXingyu Zhou, Leheng Zhang, Xiaorui Zhao, Keze Wang 等CVPR 2024 · 被引用 20 次
相关 Paper
- APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUsYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao 等ACL 2025
- LongVILA: Scaling Long-Context Visual Language Models for Long VideosYukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu 等ICLR 2025 · 被引用 1 次
- LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language ModelsTzu-Tao Chang, Shivaram VenkataramanICML 2025
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM InferenceSamir Khaki, Junxian Guo, Jiaming Tang, Shang Yang 等ICCV 2025 · 被引用 3 次
- Free-Moref: Instantly Multiplexing Context Perception Capabilities of Video-Mllms Within Single InferenceKuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang 等ICCV 2025
