NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
Qichao Wang, Ziqiao Meng, Wenqian Cui, Yifei Zhang, Pengcheng Wu, Bingzhe Wu, Irwin King, Liang Chen, Peilin Zhao
Abstract
Inspired by the impressive capabilities of GPT-4o, there is growing interest in enabling speech language models (SLMs) to engage in natural, fluid spoken interactions with humans. Recent advancements have led to the development of several SLMs that demonstrate promising results in this area. However, current approaches have yet to fully exploit dual-channel speech data, which inherently captures the structure and dynamics of human conversation. In this work, we systematically explore the use of dual-channel speech data in the context of modern large language models, and introduce a novel generative modeling paradigm-Next-Token-Pair Prediction (NTPP)-to enable speaker-independent dual-channel spoken dialogue learning using decoder-only architectures for the first time. We evaluate our approach on standard benchmarks, and empirical results show that our proposed method, NTPP, significantly improves the conversational abilities of SLMs in terms of turntaking prediction, response coherence, and naturalness. Moreover, compared to existing methods, NTPP achieves substantially lower inference latency, highlighting its practical efficiency for realtime applications. Demo and code can be found at https://audio-3059.pages.dev .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 917ad761-4526-4ce9-8b9f-eb3f30e83cd9Cited by top-tier papers6
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language ModelsChung-Ming Chien, Manu Orsini, Eugene Kharitonov, Neil Zeghidour et al.ICML 2026 · 7 citations
- Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue ModelsYifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen et al.ACL 2026 · 7 citations
- -Voice: Benchmarking Full-Duplex Voice Agents on Real-World DomainsSoham Ray, Keshav Dhandhania, Victor Barres, Karthik NarasimhanICML 2026
- Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean DatasetsWei Liu, Zhongyu Niu, Lang Gao, Zhiying Deng et al.ICML 2025
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng et al.ACL 2025
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan et al.NeurIPS 2024
- OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen et al.ACL 2025 · 51 citations
- VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality GenerationYuhao Wang, Heyang Liu, Ziyang Cheng, Ronghua Wu et al.EMNLP 2025 · 3 citations
- Language Model Can Listen While SpeakingZiyang Ma, Yakun Song, Chenpeng Du, Jian Cong et al.AAAI 2025 · 58 citations
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsYuanyuan Wang, Dongchao Yang, Yiwen Shao, Hangting Chen et al.AAAI 2026 · 3 citations
