RepCodec: A Speech Representation Codec for Speech Tokenization
Zhichao Huang, Chutong Meng, Tom Ko
摘要
With recent rapid growth of large language models (LLMs), discrete speech tokenization has played an important role for injecting speech into LLMs. However, this discretization gives rise to a loss of information, consequently impairing overall performance. To improve the performance of these discrete speech tokens, we present RepCodec, a novel speech representation codec for semantic speech tokenization. In contrast to audio codecs which reconstruct the raw audio, RepCodec learns a vector quantization codebook through reconstructing speech representations from speech encoders like HuBERT or data2vec. Together, the speech encoder, the codec encoder and the vector quantization codebook form a pipeline for converting speech waveforms into semantic tokens. The extensive experiments illustrate that RepCodec, by virtue of its enhanced information retention capacity, significantly outperforms the widely used k-means clustering approach in both speech understanding and generation. Furthermore, this superiority extends across various speech encoders and languages, affirming the robustness of RepCodec. We believe our method can facilitate large language modeling research on speech processing. Our code and models are released at https://github.com/mct10/AudioDec_ct .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- FocalCodec: Low-Bitrate Speech Coding via Focal Modulation NetworksLuca Della Libera, Francesco Paissan, Cem Subakan, Mirco RavanelliNeurIPS 2025 · 被引用 35 次
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingRui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang 等AAAI 2026 · 被引用 4 次
- SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language ModelsLinqin Wang, Yaping Liu, Zhengtao Yu, Shengxiang Gao 等AAAI 2025 · 被引用 3 次
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsYuanyuan Wang, Dongchao Yang, Yiwen Shao, Hangting Chen 等AAAI 2026 · 被引用 3 次
- Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNetsChenlin Liu, Minghui Fang, Patrick Zhang, Wei Zhou 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper6
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- PolyVoice: Language Models for Speech to Speech TranslationQianqian Dong, Zhiying Huang, Qi Tian, Chen Xu 等ICLR 2024 · 被引用 32 次
相关 Paper
- SpeechTokenizer: Unified Speech Tokenizer for Speech Language ModelsXin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou 等ICLR 2024 · 被引用 126 次
- Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelZhen Ye, Peiwen Sun, Jiahe Lei, Hongzhan Lin 等AAAI 2025 · 被引用 89 次
- Language-Codec: Bridging Discrete Codec Representations and Speech Language ModelsShengpeng Ji, Minghui Fang, Jialong Zuo, Ziyue Jiang 等ACL 2025
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream QuantizationWenxi Chen, Ruiqi Yan, Yushen Chen, Zhikang Niu 等ACL 2026 · 被引用 10 次
- CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-SpeechJaehyeon Kim, Keon Lee, Seungjun Chung, Jaewoong ChoICLR 2024 · 被引用 67 次
