SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, Xipeng Qiu
Abstract
Current speech large language models build upon discrete speech representations, which can be categorized into semantic tokens and acoustic tokens. However, existing speech tokens are not specifically designed for speech language modeling. To assess the suitability of speech tokens for building speech language models, we established the first benchmark, SLMTokBench. Our results indicate that neither semantic nor acoustic tokens are ideal for this purpose. Therefore, we propose SpeechTokenizer, a unified speech tokenizer for speech large language models. SpeechTokenizer adopts the Encoder-Decoder architecture with residual vector quantization (RVQ). Unifying semantic and acoustic tokens, SpeechTokenizer disentangles different aspects of speech information hierarchically across different RVQ layers. Furthermore, We construct a Unified Speech Language Model (USLM) leveraging SpeechTokenizer. Experiments show that SpeechTokenizer performs comparably to EnCodec in speech reconstruction and demonstrates strong performance on the SLMTokBench benchmark. Also, USLM outperforms VALL-E in zero-shot Text-to-Speech tasks. Code and models are available at https://github.com/ZhangXInFD/SpeechTokenizer/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df19bc02-ef87-4e3e-a759-56bec71f0916Cited by top-tier papers21
- TransVIP: Speech to Speech Translation System with Voice and Isochrony PreservationChenyang Le, Yao Qian, Dongmei Wang, Long Zhou et al.NeurIPS 2024 · 25 citations
- Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyTianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang et al.EMNLP 2025 · 10 citations
- Towards True Speech-to-Speech Models Without Text GuidanceXingjian Zhao, Zhe Xu, Luozhijie Jin, Yang Wang et al.ICLR 2026 · 8 citations
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech RepresentationsChao-Hong Tan, Qian Chen, Wen Wang, Chong Deng et al.ICLR 2026 · 8 citations
- Simultaneous Speech-to-Speech Translation Without Aligned DataTom Labiausse, Romain Fabre, Yannick Estève, Alexandre Défossez et al.ICML 2026 · 5 citations
Builds on4
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson et al.ICML 2020 · 210 citations
Related papers
- Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelZhen Ye, Peiwen Sun, Jiahe Lei, Hongzhan Lin et al.AAAI 2025 · 89 citations
- RepCodec: A Speech Representation Codec for Speech TokenizationZhichao Huang, Chutong Meng, Tom KoACL 2024 · 16 citations
- SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language ModelsLinqin Wang, Yaping Liu, Zhengtao Yu, Shengxiang Gao et al.AAAI 2025 · 3 citations
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingShengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen et al.ICLR 2025
- Language-Codec: Bridging Discrete Codec Representations and Speech Language ModelsShengpeng Ji, Minghui Fang, Jialong Zuo, Ziyue Jiang et al.ACL 2025
