What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
Xiaoran Fan, Zhichao Sun, Yangfan Gao, Jingfei Xiong, Hang Yan, Yifei Cao, Jiajun Sun, Shuo Li, Zhihao Zhang, Zhiheng Xi, Yuhao Zhou, Senjie Jin
摘要
Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective crossmodal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semidecoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This results in up to 12× faster decoding and a substantial reduction in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce Ro-leTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz 等ICML 2024 · 被引用 286 次
相关 Paper
- SpeechTokenizer: Unified Speech Tokenizer for Speech Language ModelsXin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou 等ICLR 2024 · 被引用 126 次
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao 等NeurIPS 2025 · 被引用 6 次
- Towards True Speech-to-Speech Models Without Text GuidanceXingjian Zhao, Zhe Xu, Luozhijie Jin, Yang Wang 等ICLR 2026 · 被引用 8 次
- VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality GenerationYuhao Wang, Heyang Liu, Ziyang Cheng, Ronghua Wu 等EMNLP 2025 · 被引用 3 次
- TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingLiang-Hsuan Tseng, Yi-Chang Chen, Kuan Yi Lee, Da-shan Shiu 等ICLR 2026 · 被引用 26 次
