Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Joon-Young Lee, Seon Joo Kim, Minho Shim
摘要
Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insight is to exploit local spatial and temporal redundancy in video data which has been overlooked in prior work. STTM first transforms each frame into multi-granular spatial tokens using a coarse-to-fine search over a quadtree structure, then performs directed pairwise merging across the temporal dimension. This decomposed merging approach outperforms existing token reduction methods across six video QA benchmarks. Notably, STTM achieves a 2 speed-up with only a 0.5% accuracy drop under a 50% token budget, and a 3 speed-up with just a 2% drop under a 30% budget. Moreover, STTM is query-agnostic, allowing KV cache reuse across different questions for the same video. The project page is available at https://www.jshyun.me/projects/sttm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token MergingZiyang Fan, Keyu Chen, Ruilong Xing, Yulin Li 等ICLR 2026 · 被引用 15 次
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand 等ICLR 2026 · 被引用 11 次
- HTTM: Head-wise Temporal Token Merging for Faster VGGTWeitian Wang, Lukas Meiner, Shubham Rai, Cecilia De la Parra 等CVPR 2026 · 被引用 7 次
- FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video UnderstandingJanghoon Cho, Jungsoo Lee, Munawar Hayat, Kyuwoong Hwang 等ICLR 2026 · 被引用 6 次
- VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-VideosWenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper27
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
相关 Paper
- MeToM: Metadata-Guided Token Merging for Efficient Video LLMsZhuojie Wu, Shijie Wang, Xin YuCVPR 2026
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai 等CVPR 2026 · 被引用 7 次
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You 等NeurIPS 2025 · 被引用 72 次
- DyCoke: Dynamic Compression of Tokens for Fast Video Large Language ModelsKeda Tao, Can Qin, Haoxuan You, Yang Sui 等CVPR 2025
- SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token PruningYicheng Ji, Jun Zhang, Heming Xia, Jinpeng Chen 等EMNLP 2025 · 被引用 1 次
