MeToM: Metadata-Guided Token Merging for Efficient Video LLMs
Zhuojie Wu, Shijie Wang, Xin Yu
摘要
Video Large Language Models (VLLMs) encounter significant computational challenges due to the large volume of visual tokens generated from multiple frames.Existing visual token pruning methods fail to account for the uneven spatiotemporal information density, thus squandering scarce token budgets on regions with low information density.In this paper, we propose a training-free Metadata-guided Token Merging framework (MeToM) that leverages intrinsic video metadata to adaptively allocate budgets and merge visual tokens based on content complexity.Specifically, MeToM exploits residual from the metadata as spatial information density cues.It merges less informative regions during tokenization, avoiding redundant encoding and improving the efficiency of the visual encoder.Additionally, MeToM captures temporal variations in information density by utilizing the average Group of Pictures (GoP) size to represent scene complexity.This mechanism enables dynamic per-frame token allocation that adaptively adjusts token budgets across time, assigning more tokens to content-complex frames and fewer to simple ones.Finally, inside the LLM, we merge low-contribution visual tokens via multi-layer attention to compact the prefill FLOPs and visual KV cache.Extensive experimental results demonstrate that MeToM outperforms the prior SoTA counterparts, achieving inference speedup against the baseline VLLM, while still improving the performance, without training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 被引用 99 次
相关 Paper
- MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMsJunpeng Ma, Qizhe Zhang, Ming Lu, Zhibin Wang 等AAAI 2026
- METok: Multi-Stage Event-based Token Compression for Efficient Long Video UnderstandingMengyue Wang, Shuo Chen, Kristian Kersting, Volker Tresp 等EMNLP 2025 · 被引用 5 次
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You 等NeurIPS 2025 · 被引用 72 次
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMsJeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim 等ICCV 2025 · 被引用 2 次
- Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language ModelsJinlong Li, Liyuan Jiang, Haonan Zhang, Nicu SebeCVPR 2026 · 被引用 5 次
