OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
Qinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, Chao-Hong Tan, Zhihao Du, Shiliang Zhang
摘要
Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a significant challenge, especially considering human conversation dynamics such as interruptions, backchannels, and overlapping speech. In this paper, we introduce a novel End-to-End GPT-based model OmniFlatten for full-duplex conversation, capable of effectively modeling the complex behaviors inherent to natural conversations with low latency. To achieve full-duplex conversation capabilities, we propose a multi-stage post-training scheme that progressively adapts a text large language model (LLM) backbone into a speech-text dialogue LLM, capable of generating text and speech in real time, without modifying the architecture of the backbone LLM. The training process comprises three stages: modality alignment, half-duplex dialogue learning, and full-duplex dialogue learning. In all training stages, we standardize the data using a flattening operation, which enables unifying the training methods and the GPT backbone across different modalities and tasks. Our approach offers a simple modeling technique and a promising research direction for developing efficient and natural end-to-end full-duplex spoken dialogue systems. Audio samples of dialogues generated by OmniFlatten can be found at this web site (https://omniflatten.github.io/).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen 等NeurIPS 2025 · 被引用 43 次
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech RepresentationsChao-Hong Tan, Qian Chen, Wen Wang, Chong Deng 等ICLR 2026 · 被引用 8 次
- SageLM: A Multi-aspect and Explainable Large Language Model for Speech JudgementYuan Ge, Junxiang Zhang, Xiaoqian Liu, Bei Li 等AAAI 2026 · 被引用 5 次
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech SynthesisRun Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu 等NeurIPS 2025 · 被引用 5 次
- End-to-end Listen, Look, Speak and ActSiyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
- Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsNing Ding, Yulin Chen, Bokai Xu, Yujia Qin 等EMNLP 2023 · 被引用 95 次
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel 等ICLR 2023 · 被引用 87 次
相关 Paper
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsBandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong 等EMNLP 2024 · 被引用 8 次
- Language Model Can Listen While SpeakingZiyang Ma, Yakun Song, Chenpeng Du, Jian Cong 等AAAI 2025 · 被引用 58 次
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan 等NeurIPS 2024
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMXiong Wang, Yangze Li, Chaoyou Fu, Yike Zhang 等ICML 2025
- NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair PredictionQichao Wang, Ziqiao Meng, Wenqian Cui, Yifei Zhang 等ICML 2025
