WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Shengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Xize Cheng, Zehan Wang, Ruiqi Li, Ziang Zhang, Xiaoda Yang
Abstract
Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous state-of-the-art (SOTA) acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective reconstruction quality. Despite the reduced number of tokens, WavTokenizer achieves SOTA reconstruction quality with outstanding UTMOS scores and also inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extending contextual windows, improving attention networks, and introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conduct extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibits competitive to superior performance across various objective and subjective metrics compared to SOTA models. We also evaluate WavTokenizer on semantic representation, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdd25011-45ea-4b62-82ea-2176dcd0b034Cited by top-tier papers43
- Multimodal Latent Language Modeling with Next-Token DiffusionYutao Sun, Hangbo Bao, Wenhui Wang, Zhiliang Peng et al.ICML 2026 · 54 citations
- LeVo: High-Quality Song Generation with Multi-Preference AlignmentShun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang et al.NeurIPS 2025 · 43 citations
- FocalCodec: Low-Bitrate Speech Coding via Focal Modulation NetworksLuca Della Libera, Francesco Paissan, Cem Subakan, Mirco RavanelliNeurIPS 2025 · 35 citations
- EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic IntegrationMinjie Hong, Yan Xia, Zehan Wang, Jieming Zhu et al.WWW 2025 · 30 citations
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame RatesJiaqi Li, Yao Qian, Yuxuan Hu, leying zhang et al.ICLR 2026 · 27 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
Related papers
- ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language ModelingDongchao Yang, Songxiang Liu, Haohan Guo, Jiankun Zhao et al.ICML 2025
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language ModelingYuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang et al.NeurIPS 2025 · 22 citations
- SpeechTokenizer: Unified Speech Tokenizer for Speech Language ModelsXin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou et al.ICLR 2024 · 126 citations
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech CodecsYitian Gong, Luozhijie Jin, Kuangwei Chen, Dong Zhang et al.ACL 2026 · 35 citations
