Lune

ICLR2025顶会

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Shengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Xize Cheng, Zehan Wang, Ruiqi Li, Ziang Zhang, Xiaoda Yang

2025年份
43顶会引用

摘要

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous state-of-the-art (SOTA) acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective reconstruction quality. Despite the reduced number of tokens, WavTokenizer achieves SOTA reconstruction quality with outstanding UTMOS scores and also inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extending contextual windows, improving attention networks, and introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conduct extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibits competitive to superior performance across various objective and subjective metrics compared to SOTA models. We also evaluate WavTokenizer on semantic representation, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext fdd25011-45ea-4b62-82ea-2176dcd0b034

引用它的顶会 Paper43

问问它们各自怎么用它

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖