Scaling Transformers for Low-Bitrate High-Quality Speech Coding
Julian D. Parker, Anton Smirnov, Jordi Pons, CJ Carr, Zack Zukowski, Zach Evans, Xubo Liu
Abstract
The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with strong inductive biases. In this work we show that by scaling a transformer architecture with large parameter count to this problem, and applying a flexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible to reach state-of-the-art speech quality at extremely low bit-rates of 400 or 700 bits-per-second. The trained models strongly out-perform existing baselines in both objective and subjective tests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c44f24eb-5d92-4081-8456-99dd9dae2427Cited by top-tier papers20
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang et al.ICLR 2026 · 112 citations
- FocalCodec: Low-Bitrate Speech Coding via Focal Modulation NetworksLuca Della Libera, Francesco Paissan, Cem Subakan, Mirco RavanelliNeurIPS 2025 · 35 citations
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame RatesJiaqi Li, Yao Qian, Yuxuan Hu, leying zhang et al.ICLR 2026 · 27 citations
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language ModelingYuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang et al.NeurIPS 2025 · 22 citations
- Scaling Speech Tokenizers with Diffusion AutoencodersYuancheng Wang, Zhenyu Tang, Yun Wang, Arthur Hinsvark et al.ICLR 2026 · 5 citations
Builds on9
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
Related papers
- Scaling Transformers for End-to-End Discrete Audio TokenizationYitian Gong, Kuangwei Chen, Zhaoye Fei, Xiaogui Yang et al.ICML 2026
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream QuantizationWenxi Chen, Ruiqi Yan, Yushen Chen, Zhikang Niu et al.ACL 2026 · 10 citations
- ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized TransformersYuzhe Gu, Enmao DiaoEMNLP 2024 · 5 citations
- Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationWenrui Liu, Qian Chen, Wen Wang, Guanrou Yang et al.ACM MM 2025
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingShengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen et al.ICLR 2025
