Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
Siddhant Arora, Haidar Khan, Kai Sun, Xin Dong, Sajal Choudhary, Seungwhan Moon, Xinyuan Zhang, Adithya Sagar, Surya Appini, Kaushik Patnaik, Sanat Sharma, Shinji Watanabe
Abstract
End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with significantly lower latency. However, these systems remain prone to hallucinations due to limited factual grounding. While text-based dialogue systems address this challenge by integrating tools such as web search and knowledge graph APIs, we introduce the first approach to extend tool use directly into speechin speech-out systems. A key challenge is that tool integration substantially increases response latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Streaming RAG), a novel framework that reduces user-perceived latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. Specifically, we develop a post-training pipeline that teaches the model when to issue tool calls during ongoing speech and how to generate spoken summaries that fuse audio queries with retrieved text results, thereby improving both accuracy and responsiveness. To evaluate our approach, we construct AudioCRAG, a benchmark created by converting queries from the publicly available CRAG dataset into speech form. Experimental results demonstrate that our streaming RAG approach increases QA accuracy by up to 200% relative (from 11.1% to 34.2% absolute) and further enhances user experience by reducing tool use latency by 20%. Importantly, our streaming RAG approach is modality-agnostic and can be applied equally to typed input, paving the way for more agentic, real-time AI assistants.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7aa7ae89-653c-4b2e-87a4-f5fe46836632Cited by top-tier papers6
- Shanks: Simultaneous Hearing and Thinking for Spoken Language ModelsCheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin et al.ACL 2026 · 14 citations
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language ModelsChung-Ming Chien, Manu Orsini, Eugene Kharitonov, Neil Zeghidour et al.ICML 2026 · 7 citations
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion ForcingWilliam Chen, Prem Seetharaman, Rithesh Kumar, Oriol Nieto et al.ICML 2026 · 7 citations
- VoxMind: An End-to-End Agentic Spoken Dialogue SystemTianle Liang, Yifu Chen, Shengpeng Ji, Yijun Chen et al.ACL 2026 · 1 citation
- ProactiveLLM: Learning Active Interaction for Streaming Large Language ModelsJunlong Tong, Yao Zhang, Anhao Zhao, Yingqi Fan et al.ICML 2026
Builds on7
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin et al.NeurIPS 2025 · 164 citations
- Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex ModelsXinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han et al.EMNLP 2024 · 4 citations
Related papers
- WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue ModelsYifu Chen, Shengpeng Ji, Haoxiao Wang, Ziqing Wang et al.ACL 2025
- StreamRAG: Enhancing Real-Time Video Understanding with Retrieval AugmentationJunlin Xie, Quanlong Zheng, Ruifei Zhang, Kuo Wang et al.CVPR 2026
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du et al.SOSP 2025 · 3 citations
- Omnia: Efficient RAG Serving through Speculative SchedulingRongtian Fu, Shigang Li, Youxuan Xu, Tong Wu et al.HPDC 2026
- Empowering GraphRAG with Knowledge Filtering and IntegrationKai Guo, Harry Shomer, Shenglai Zeng, Haoyu Han et al.EMNLP 2025 · 2 citations
