StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation
Junlin Xie, Quanlong Zheng, Ruifei Zhang, Kuo Wang, Yanhao Zhang, Jinguo Luo, Haonan Lu, Xiang Wan, Guanbin Li
Abstract
Retrieval-Augmented Generation (RAG) has shown considerable promise in offline video comprehension; however, its application to streaming video remains relatively unexplored. Streaming video introduces unique challenges, such as continuous data influx, temporal sensitivity, and stringent latency requirements. Key obstacles in deploying RAG for streaming video include: (1) the necessity for adaptive semantic segmentation to enable real-time boundary detection; (2) the challenge of balancing latency and accuracy in knowledge extraction; and (3) the complexity of handling queries with varying degrees of temporal sensitivity. To address these issues, we present StreamRAG, an innovative framework designed for streaming video question answering. StreamRAG integrates: (1) a Stream Event Segmentation (SES) module that divides video streams into semantically coherent events; (2) a knowledge extraction accelerator that minimizes captioning latency by reusing previously processed tokens; and (3) a query-aware dynamic knowledge injection module that optimizes retrieval based on the temporal sensitivity of queries and similarity scoring. Experimental results demonstrate that StreamRAG significantly enhances the efficiency of real-time video comprehension while maintaining a balance between accuracy and responsiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb30250c-4e7f-4458-ac6d-d8632772a2bfBuilds on20
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin et al.NeurIPS 2025 · 164 citations
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantHaibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu et al.NeurIPS 2025 · 63 citations
- ViSpeak: Visual Instruction Feedback in Streaming VideosShenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng et al.ICCV 2025 · 42 citations
Related papers
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li et al.AAAI 2026 · 1 citation
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video UnderstandingZhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai et al.NeurIPS 2025 · 19 citations
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and CompressionYilong Chen, Xiang Bai, Zhibin Wang, Chengyu Bai et al.AAAI 2026 · 1 citation
- QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive ResponseKairui Zhang, Zhenyu Yang, Bing Wang, Shengsheng Qian et al.ICLR 2026
- MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented GenerationYubo Wang, Haoyang Li, Lei ChenVLDB 2026
