StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Xin Ding, Hao Wu, Yifan Yang, Shiqi Jiang, Qianxi Zhang, Donglin Bai, Zhibo Chen, Ting Cao
摘要
With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce STREAMMIND, a video LLM framework that achieves ultra-FPS streaming video processing (100 fps on a single A100) and enables proactive, always-on responses in real time, without explicit user intervention.
To solve the key challenge of the contradiction between linear video streaming speed and quadratic transformer computation cost, we propose a novel perception-cognition interleaving paradigm named "event-gated LLM invocation", in contrast to the existing per-time-step LLM invocation. By introducing a Cognition Gate network between the video encoder and the LLM, LLM is only invoked when relevant events occur. To realize the event feature extraction with constant cost, we propose Event-Preserving Feature Extractor (EPFE) based on state-space method, generating a single perception token for spatiotemporal features. These techniques enable the video LLM with full-FPS perception and real-time cognition response. Experiments on Ego4D and SoccerNet streaming tasks, as well as standard offline benchmarks, demonstrate stateof-the-art performance in both model capability and realtime efficiency, paving the way for ultra-high-FPS applications, such as Game AI and interactive media. The code and data is available at https://aka.ms/StreamMind. * This work was done during Xin Ding's internship at Microsoft Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li 等NeurIPS 2025 · 被引用 79 次
- StreamingTOM: Streaming Token Compression for Efficient Video UnderstandingXueyi Chen, Keda Tao, Kele Shao, Huan WangCVPR 2026 · 被引用 46 次
- LiveStar: Live Streaming Assistant for Real-World Online Video UnderstandingZhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang 等NeurIPS 2025 · 被引用 26 次
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh, Reza Pourreza 等NeurIPS 2025 · 被引用 2 次
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and CompressionYilong Chen, Xiang Bai, Zhibin Wang, Chengyu Bai 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji 等ICML 2024 · 被引用 527 次
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
相关 Paper
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and ReactionRui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang 等CVPR 2025
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeHaomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge 等ICLR 2025
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun 等CVPR 2026 · 被引用 28 次
- VideoLLM-online: Online Video Large Language Model for Streaming VideoJoya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin 等CVPR 2024
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 被引用 34 次
