Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Yiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zenghui Ding, Xianjun Yang, Yining Sun
摘要
Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving contextrich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding. Code is available at https://github.com/Jam1ezhang/DYTO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- Efficient Multi-modal Large Language Models via Progressive Consistency DistillationZichen Wen, Shaobo Wang, Yufa Zhou, Junyuan Zhang 等NeurIPS 2025 · 被引用 26 次
- AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video GenerationHaoyue Tan, Shengnan Wang, Yulin Qiao, Juncheng Zhang 等CVPR 2026 · 被引用 5 次
- video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLMGuangzhi Sun, Yixuan Li, Xiaodong Wu, Yudong Yang 等ICML 2026 · 被引用 2 次
- A Multi-Agent Perception-Action Alliance for Efficient Long Video ReasoningYichang Xu, Gaowen Liu, Ramana Rao Kompella, Tiansheng Huang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper18
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
相关 Paper
- MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMsJunpeng Ma, Qizhe Zhang, Ming Lu, Zhibin Wang 等AAAI 2026
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMsJeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim 等ICCV 2025 · 被引用 2 次
- MeToM: Metadata-Guided Token Merging for Efficient Video LLMsZhuojie Wu, Shijie Wang, Xin YuCVPR 2026
- FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token MergingZiyang Fan, Keyu Chen, Ruilong Xing, Yulin Li 等ICLR 2026 · 被引用 15 次
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 被引用 1 次
