Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation
Weidong Chen, Helen M. Meng, Xixin Wu
摘要
Current speech language models require their core component, the speech codec, to discretize continuous speech signals into tokens that not only capture high-level cues for autoregressive modeling but also preserve sufficient acoustic details for perceptual quality. To address this need, we propose Gogo, a group-wise granularity-ordered codec that quantizes each group of frames into tokens arranged from coarse to fine, where coarse tokens encode high-level abstractions and fine tokens progressively recover low-level details. Building on the granularity-ordering property of Gogo, we introduce GogoSpeech, a two-stage speech language model that performs speech generation by first constructing a coarse speech backbone at an extremely low token rate and then enriching the backbone with fine-grained acoustic details. Considering the inherently non-uniform information distribution in speech signals, we further design a Group Relative Policy Optimization (GRPO)-trained token allocator that adaptively allocates token budgets to groups based on group-wise complexity. Experimental results demonstrate that Gogo delivers state-of-the-art reconstruction performance across most metrics at a token rate of 47. Moreover, evaluations on zero-shot text-to-speech tasks show that GogoSpeech enables efficient generation by adaptively reducing the average token rate, and attains state-of-the-art results in long-form speech generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
相关 Paper
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingShengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen 等ICLR 2025
- Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationWenrui Liu, Qian Chen, Wen Wang, Guanrou Yang 等ACM MM 2025
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingRui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang 等AAAI 2026 · 被引用 4 次
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language ModelingYuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang 等NeurIPS 2025 · 被引用 22 次
- SpeechTokenizer: Unified Speech Tokenizer for Speech Language ModelsXin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou 等ICLR 2024 · 被引用 126 次
