VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation
Yuhao Wang, Heyang Liu, Ziyang Cheng, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang
Abstract
Speech large language models (LLMs) have emerged as a prominent research focus in speech processing. In this work, we introduce VocalNet, a series of high-performance speech LLMs featuring a scalable and model-agnostic training framework as well as a novel multitoken prediction (MTP) paradigm for speech generation. We first propose an efficient twostage training framework that enables LLMs to acquire real-time speech interaction capabilities. Through extensive experiments on various training configurations, we ensure both simplicity and effectiveness in the training strategy. Furthermore, inspired by advances in language modeling, we introduce MTP into the domain of speech LLMs-an alternative to traditional next-token prediction (NTP)-which enables the model to predict multiple future tokens at each step. Through systematic analysis and improved implementation, we show that MTP not only accelerates inference speed but also significantly enhances speech quality. Experimental results demonstrate that VocalNet achieves performance comparable to state-ofthe-art Omni LLMs while outperforming existing open-source speech LLMs, despite using limited training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfcd3339-9db4-47db-911e-2e72f5a8665aBuilds on9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
Related papers
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang et al.NeurIPS 2025 · 16 citations
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao et al.NeurIPS 2025 · 6 citations
- What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyXiaoran Fan, Zhichao Sun, Yangfan Gao, Jingfei Xiong et al.AAAI 2026
- NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair PredictionQichao Wang, Ziqiao Meng, Wenqian Cui, Yifei Zhang et al.ICML 2025
- LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisQingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang et al.ACL 2025
