SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inference
Haojuan Li, Ruohan Tang, Dongzhou Cheng, Zongpu Zhang, Jian Li, Jiaqi Wang
摘要
Video Large Language Models (VideoLLMs) face a fundamental performance bottleneck: the token explosion intrinsic to video inputs. The resulting O(N 2 ) prefill cost makes conventional Transformer inference prohibitively expensive at scale. Existing attempts fall into a hard accuracylatency dilemma: naive sparsification risks losing essential temporal-spatial context, whereas naive parallelization introduces substantial communication and memory overhead. To overcome this impasse, we propose SegMo, a unified framework based on algorithm-system co-design that jointly optimizes what to compute (sparsification) and how to compute it (parallelism). Driven by the empirical discovery of Local Cohesion in VideoLLM attention, SegMo integrates two key components: (1) Content-Aware Sparsification (CAS): A lightweight, hierarchical algorithm that first employs Query Relevance for scene-level assessment, and then uses Temporal Redundancy for intra-scene static redundancy pruning, to generate a precise, non-uniform computation load, ensuring accuracy. (2) Locally-Cohesive Segment Parallelism (LSP): A novel paradigm that leverages attention locality to partition the video at scene boundaries, using a lightweight Global Context Injection mechanism to replace the massive communication and memory overheads of global attention. SegMo was validated across LVBench, LongVideoBench, and Video-MME. Our CAS module improved accuracy by up to 12.00%. When integrated with LSP, the full system (CAS + LSP) achieved a peak prefill acceleration of 3.55×, while still maintaining a significant accuracy gain of up to 8.31%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
相关 Paper
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai 等CVPR 2026 · 被引用 7 次
- MeToM: Metadata-Guided Token Merging for Efficient Video LLMsZhuojie Wu, Shijie Wang, Xin YuCVPR 2026
- Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language ModelsZhenyu Li, Zuchao Li, Ping Wang, Lefei Zhang 等ACL 2026
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM InferenceSamir Khaki, Junxian Guo, Jiaming Tang, Shang Yang 等ICCV 2025 · 被引用 3 次
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You 等NeurIPS 2025 · 被引用 72 次
