Lune

NeurIPS2025顶会

Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs

Evangelos Dorovatas, Soroush Seifi, Gunshi Gupta, Rahaf Aljundi

2025年份
8被引次数
1顶会引用

摘要

Video Large Language Models (Video-LLMs) excel at understanding videos incontext, provided they have full access to the video when answering queries. However, these models face challenges in streaming scenarios where hour-long videos must be processed online, and questions need timely responses. In this work, we propose a training-free approach compatible with standard Video-LLMs, leveraging three key concepts: 1) LLM-informed selection of visual tokens to identify those that the LLM has attended to and contributed to its understanding of each short clip. Our attention-based selection allows us to discard up to ∼ 95% of unimportant visual tokens with minimal performance loss; 2) Recurrent processing of past selected tokens to generate temporally coherent understanding of each processed clip; 3) Caption-based question answering for lightweight and accurate responses. Our method achieves state-of-the-art performance on streaming video benchmarks, striking a balance between efficiency and effectiveness.

Recent efforts have focused on compressing the information from short video clips, moving away from the brute-force approach of processing the entire video in a single pass, and thus extending model capabilities to handle long video understanding. These approaches include methods that either compress only the visual information (before the LLM) encoded by a vision encoder [38,12,10], or store only textual descriptions of short clips [1], and retrieve only those relevant to the input *First two authors provide contracted services for Toyota.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖