Lune

CVPR2026顶会

StreamingTOM: Streaming Token Compression for Efficient Video Understanding

Xueyi Chen, Keda Tao, Kele Shao, Huan Wang

2026年份
46被引次数
9顶会引用

摘要

Unlike offline processing, streaming video vision-language models face two fundamental constraints: causality and accumulation. Causality prevents access to future frames that offline methods exploit, while accumulation causes tokens to grow unbounded, creating efficiency bottlenecks. However, existing approaches only regulate post-LLM kv-cache, leaving costly pre-LLM prefill unchanged. We introduce StreamingTOM, a training-free, plug-and-play two-stage framework that addresses both pre-LLM and post-LLM bottlenecks. Causal Temporal Reduction imposes a fixed per-frame budget and selects tokens based on adjacent-frame changes and token saliency, drastically reducing per-frame prefill cost by processing only a compact subset of visual tokens, ensuring predictable latency. Online Quantized Memory stores tokens in 4-bit format, retrieves relevant groups on demand, and dequantizes them, keeping the active kv-cache bounded regardless of stream length. Experiments demonstrate our method achieves 15.7×15.7\times kv-cache compression ratio; compared to prior SOTA (LiveVLM), it delivers 1.2×1.2\times lower peak memory and 2×2\times faster TTFT. StreamingTOM achieves state-of-the-art accuracy among training-free methods with an average of 63.8%63.8\% on offline benchmarks and 55.8%55.8\% accuracy and 3.73.7 score on RVS. These results demonstrate that real-time streaming video understanding with bounded active memory is achievable without model retraining.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 5d5a1814-9fda-48c5-8b94-f1a453a91316

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper35

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖