MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
Chung-Ming Chien, Manu Orsini, Eugene Kharitonov, Neil Zeghidour, Karen Livescu, Alexandre Défossez
摘要
Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time interactivity, including handling of pauses, interruptions, and backchannels. However, improving their factuality remains an open challenge. While scaling the model size could address this gap, it would make real-time inference prohibitively expensive. In this work, we propose Moshi-RAG, a modular approach that combines a compact full-duplex interface with selective retrieval to access more powerful knowledge sources. Our asynchronous framework enables the model to identify knowledge-demanding queries and ground its responses in external information. By leveraging the natural temporal gap between response onset and the delivery of core information, the retrieval process can be completed while maintaining a natural conversation flow. With this approach, Moshi-RAG achieves factuality comparable to the best publicly released non-duplex speech language models while preserving the interactivity inherent to full-duplex systems. Moreover, our flexible design supports plug-and-play retrieval methods without retraining and demonstrates strong performance on out-of-domain mathematical reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 等CVPR 2022 · 被引用 184 次
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMEliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar 等ICLR 2024 · 被引用 95 次
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen 等NeurIPS 2025 · 被引用 43 次
相关 Paper
- Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool UsageSiddhant Arora, Haidar Khan, Kai Sun, Xin Dong 等ICML 2026 · 被引用 22 次
- Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMsZhenyu Liu, Xuanyu Zhang, Yunxin Li, Qixun Teng 等ACL 2026
- RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware ReasoningYu Wang, Shiwan Zhao, Zhihu Wang, Ming Fan 等EMNLP 2025 · 被引用 3 次
- Vision-Speech Models: Teaching Speech Models to Converse about ImagesAmélie Royer, Moritz Böhle, Laurent Mazaré, Neil Zeghidour 等CVPR 2026 · 被引用 4 次
- DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question AnsweringRong Cheng, Jinyi Liu, Yan Zheng, Fei Ni 等ACL 2025
