StreamingTOM: Streaming Token Compression for Efficient Video Understanding
Xueyi Chen, Keda Tao, Kele Shao, Huan Wang
Abstract
Unlike offline processing, streaming video vision-language models face two fundamental constraints: causality and accumulation. Causality prevents access to future frames that offline methods exploit, while accumulation causes tokens to grow unbounded, creating efficiency bottlenecks. However, existing approaches only regulate post-LLM kv-cache, leaving costly pre-LLM prefill unchanged. We introduce StreamingTOM, a training-free, plug-and-play two-stage framework that addresses both pre-LLM and post-LLM bottlenecks. Causal Temporal Reduction imposes a fixed per-frame budget and selects tokens based on adjacent-frame changes and token saliency, drastically reducing per-frame prefill cost by processing only a compact subset of visual tokens, ensuring predictable latency. Online Quantized Memory stores tokens in 4-bit format, retrieves relevant groups on demand, and dequantizes them, keeping the active kv-cache bounded regardless of stream length. Experiments demonstrate our method achieves kv-cache compression ratio; compared to prior SOTA (LiveVLM), it delivers lower peak memory and faster TTFT. StreamingTOM achieves state-of-the-art accuracy among training-free methods with an average of on offline benchmarks and accuracy and score on RVS. These results demonstrate that real-time streaming video understanding with bounded active memory is achievable without model retraining.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You et al.NeurIPS 2025 · 72 citations
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language ModelsKeda Tao, Kele Shao, Bohan Yu, Weiqiang Wang et al.CVPR 2026 · 32 citations
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang et al.ICLR 2026 · 26 citations
- HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video UnderstandingHaowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng et al.ACL 2026 · 17 citations
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal UnderstandingXin Jin, Siyuan Li, Siyong Jian, Kai Yu et al.ICLR 2026 · 14 citations
Builds on35
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You et al.NeurIPS 2025 · 72 citations
Related papers
- Decouple and Cache: KV Cache Construction for Streaming Video UnderstandingZhanzhong Pang, Dibyadip Chatterjee, Fadime Sener, Angela YaoICML 2026 · 1 citation
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video UnderstandingMinsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung ChangNeurIPS 2025 · 62 citations
- FluxMem: Adaptive Hierarchical Memory for Streaming Video UnderstandingYiweng Xie, Bo He, Junke Wang, Xiangyu Zheng et al.CVPR 2026 · 25 citations
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and CompressionYilong Chen, Xiang Bai, Zhibin Wang, Chengyu Bai et al.AAAI 2026 · 1 citation
- Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMsEvangelos Dorovatas, Soroush Seifi, Gunshi Gupta, Rahaf AljundiNeurIPS 2025 · 8 citations
