Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
Bandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong, Shyamnath Gollakota
Abstract
Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking of interruption or silence events. Human dialogue, by contrast, is "full-duplex" allowing for rich synchronicity in the form of quick and dynamic turn-taking, overlapping speech, and backchanneling. Technically, the challenge of achieving full-duplex dialogue with LLMs lies in modeling synchrony as pre-trained LLMs do not have a sense of "time". To bridge this gap, we propose Synchronous LLMs for fullduplex spoken dialogue modeling. We design a novel mechanism to integrate time information into Llama3-8b so that they run synchronously with the real-world clock. We also introduce a training recipe that uses 212k hours of synthetic spoken dialogue data generated from text dialogue data to create a model that generates meaningful and natural spoken dialogue, with just 2k hours of real-world spoken dialogue data. Synchronous LLMs outperform state-of-the-art in dialogue meaningfulness while maintaining naturalness. Finally, we demonstrate the model's ability to participate in full-duplex dialogue by simulating interaction between two agents trained on different datasets, while considering Internet-scale latencies of up to 240ms. Webpage: https: //syncllm.cs.washington.edu/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8b70a32-c9f6-4d4e-a3b6-578beb8da54eCited by top-tier papers16
- OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen et al.ACL 2025 · 51 citations
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen et al.NeurIPS 2025 · 43 citations
- WearVox: An Egocentric Multichannel Voice Assistant Benchmark for WearablesZhaojiang Lin, Yong Xu, Kai Sun, Jing Zheng et al.ICLR 2026 · 11 citations
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-trainingDingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo et al.ACL 2025 · 9 citations
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language ModelsChung-Ming Chien, Manu Orsini, Eugene Kharitonov, Neil Zeghidour et al.ICML 2026 · 7 citations
Builds on4
- Knowledge-Grounded Dialogue Generation with Pre-trained Language ModelsXueliang Zhao, Wei Wu, Can Xu, Chongyang Tao et al.EMNLP 2020 · 153 citations
- Textually Pretrained Speech Language ModelsMichael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat et al.NeurIPS 2023 · 117 citations
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMEliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar et al.ICLR 2024 · 95 citations
- Text-Free Prosody-Aware Generative Spoken Language ModelingEugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi et al.ACL 2022
Related papers
- Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex ModelsXinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han et al.EMNLP 2024 · 4 citations
- Language Model Can Listen While SpeakingZiyang Ma, Yakun Song, Chenpeng Du, Jian Cong et al.AAAI 2025 · 58 citations
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan et al.NeurIPS 2024
- Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken ConversationsGuan-Ting Lin, Cheng-Han Chiang, Hung-yi LeeACL 2024 · 15 citations
- NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair PredictionQichao Wang, Ziqiao Meng, Wenqian Cui, Yifei Zhang et al.ICML 2025
