ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
Yuzhe Gu, Enmao Diao
摘要
Neural speech codecs aim to compress input signals into minimal bits while maintaining content quality in a low-latency manner. However, existing neural codecs often trade model complexity for reconstruction performance. These codecs primarily use convolutional blocks for feature transformation, which are not inherently suited for capturing the local redundancies in speech signals. To compensate, they require either adversarial discriminators or a large number of model parameters to enhance audio quality. In response to these challenges, we introduce the Efficient Speech Codec (ESC) 1 , a lightweight, parameter-efficient speech codec based on a cross-scale residual vector quantization scheme and transformers. Our model employs mirrored hierarchical window transformer blocks and performs step-wise decoding from coarse-tofine feature representations. To enhance bitrate efficiency, we propose a novel combination of vector quantization techniques along with a pre-training paradigm. Extensive experiments demonstrate that ESC can achieve high-fidelity speech reconstruction with significantly lower model complexity, making it a promising alternative to existing convolutional audio codecs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
相关 Paper
- Scaling Transformers for Low-Bitrate High-Quality Speech CodingJulian D. Parker, Anton Smirnov, Jordi Pons, CJ Carr 等ICLR 2025
- Two-dimensional quantization for geometry-aware audio codingTal Shuster, Eliya NachmaniICML 2026
- Self-Guidance: Enhancing Neural Codecs via Decoder Manifold AlignmentXiang Li, Yixuan Zhou, Xie, Zhiyong Wu 等ICML 2026
- Language-Codec: Bridging Discrete Codec Representations and Speech Language ModelsShengpeng Ji, Minghui Fang, Jialong Zuo, Ziyue Jiang 等ACL 2025
- FocalCodec: Low-Bitrate Speech Coding via Focal Modulation NetworksLuca Della Libera, Francesco Paissan, Cem Subakan, Mirco RavanelliNeurIPS 2025 · 被引用 35 次
