Language Model Can Listen While Speaking
Ziyang Ma, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen, Yuping Wang, Yuxuan Wang, Xie Chen
Abstract
Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM), have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based conversation, lacking the ability to interact with humans in real-time spoken scenarios, for example, being interrupted when the generated content is not satisfactory. To address these limitations, we explore full duplex modeling (FDM) in interactive speech language models (iSLM), focusing on enhancing real-time interaction and, more explicitly, exploring the quintessential ability of interruption. We introduce a novel model design, namely listening-while-speaking language model (LSLM), an end-to-end system equipped with both listening and speaking channels. Our LSLM employs a token-based decoder-only TTS for speech generation and a streaming self-supervised learning (SSL) encoder for real-time audio input. LSLM fuses both channels for autoregressive generation and detects turn-taking in real time. Three fusion strategies-early fusion, middle fusion, and late fusion-are explored, with middle fusion achieving an optimal balance between speech generation and real-time interaction. Two experimental settings, command-based FDM and voice-based FDM, demonstrate LSLM's robustness to noise and sensitivity to diverse instructions. Our results highlight LSLM's capability to achieve duplex communication with minimal impact on existing systems. This study aims to advance the development of interactive speech dialogue systems, enhancing their applicability in real-world contexts 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4524458d-9fd8-4543-9f3d-2ff2fbfab08bCited by top-tier papers15
- OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen et al.ACL 2025 · 51 citations
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen et al.NeurIPS 2025 · 43 citations
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue et al.ACL 2026 · 17 citations
- Speculative End-Turn Detector for Efficient Speech Chatbot AssistantHyunjong Ok, Suho Yoo, Jaeho LeeACL 2026 · 5 citations
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMsYuhan Song, Linhao Zhang, Chuhan Wu, Aiwei Liu et al.ICLR 2026 · 5 citations
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky et al.ICLR 2024 · 247 citations
Related papers
- Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex ModelsXinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han et al.EMNLP 2024 · 4 citations
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan et al.NeurIPS 2024
- NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair PredictionQichao Wang, Ziqiao Meng, Wenqian Cui, Yifei Zhang et al.ICML 2025
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsBandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong et al.EMNLP 2024 · 8 citations
- Aligning Spoken Dialogue Models from User InteractionsAnne Wu, Laurent Mazaré, Neil Zeghidour, Alexandre DéfossezICML 2025
