StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation
Junlin Xie, Quanlong Zheng, Ruifei Zhang, Kuo Wang, Yanhao Zhang, Jinguo Luo, Haonan Lu, Xiang Wan, Guanbin Li
摘要
Retrieval-Augmented Generation (RAG) has shown considerable promise in offline video comprehension; however, its application to streaming video remains relatively unexplored. Streaming video introduces unique challenges, such as continuous data influx, temporal sensitivity, and stringent latency requirements. Key obstacles in deploying RAG for streaming video include: (1) the necessity for adaptive semantic segmentation to enable real-time boundary detection; (2) the challenge of balancing latency and accuracy in knowledge extraction; and (3) the complexity of handling queries with varying degrees of temporal sensitivity. To address these issues, we present StreamRAG, an innovative framework designed for streaming video question answering. StreamRAG integrates: (1) a Stream Event Segmentation (SES) module that divides video streams into semantically coherent events; (2) a knowledge extraction accelerator that minimizes captioning latency by reusing previously processed tokens; and (3) a query-aware dynamic knowledge injection module that optimizes retrieval based on the temporal sensitivity of queries and similarity scoring. Experimental results demonstrate that StreamRAG significantly enhances the efficiency of real-time video comprehension while maintaining a balance between accuracy and responsiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin 等NeurIPS 2025 · 被引用 164 次
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantHaibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu 等NeurIPS 2025 · 被引用 63 次
- ViSpeak: Visual Instruction Feedback in Streaming VideosShenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng 等ICCV 2025 · 被引用 42 次
相关 Paper
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li 等AAAI 2026 · 被引用 1 次
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video UnderstandingZhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai 等NeurIPS 2025 · 被引用 19 次
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and CompressionYilong Chen, Xiang Bai, Zhibin Wang, Chengyu Bai 等AAAI 2026 · 被引用 1 次
- QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive ResponseKairui Zhang, Zhenyu Yang, Bing Wang, Shengsheng Qian 等ICLR 2026
- MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented GenerationYubo Wang, Haoyang Li, Lei ChenVLDB 2026
