Lune

ACM MM2025顶会

CITR: Efficient Long Video Understanding Needs Causal Importance

Ziqi Yuan, Jun Li, Yanghao Li, Yuxiang Huang, Chi Chen, Shuo Wang, Zhinan Gou

2025年份

摘要

Long video understanding is essential for various practical applications including surveillance and film analysis. While recent Vision-Language Models (VLMs) have advanced performance in this domain, efficiency remains a key challenge, especially for hour-long videos. Existing methods commonly reduce visual tokens via compression in the vision encoder, but token count still grows linearly with video length. Alternative approaches apply importance-based token reduction in the language model, yet their non-causal design limits efficiency gains to offline, single-query settings. In this work, we emphasize the need for causal importance estimation-where a token's relevance is determined only from prior context-to enable efficient, real-time long video understanding. We propose ØurMethod, a Causal Importance-based Token Reduction framework to reduce visual token redundancy in long video understanding tasks, enabling practical memory control and enhanced computational efficiency. Experiments on both offline and streaming benchmarks show that ØurMethod reduces latency by 49% in offline multi-query scenarios and effectively controls chunked prefilling time in streaming, all within a 24GB memory footprint and with less than 1% performance drop. The code and appendix are available at https://github.com/Columbine21/CITR.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖