Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Hubert Siuzdak
摘要
Recent advancements in neural vocoding are predominantly driven by Generative Adversarial Networks (GANs) operating in the time-domain. While effective, this approach neglects the inductive bias offered by time-frequency representations, resulting in reduntant and computionally-intensive upsampling operations. Fourier-based time-frequency representation is an appealing alternative, aligning more accurately with human auditory perception, and benefitting from well-established fast algorithms for its computation. Nevertheless, direct reconstruction of complex-valued spectrograms has been historically problematic, primarily due to phase recovery issues. This study seeks to close this gap by presenting Vocos, a new model that directly generates Fourier spectral coefficients. Vocos not only matches the state-of-the-art in audio quality, as demonstrated in our evaluations, but it also substantially improves computational efficiency, achieving an order of magnitude increase in speed compared to prevailing time-domain neural vocoding approaches. The source code and model weights have been open-sourced at https://github.com/gemelo-ai/vocos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar 等NeurIPS 2025 · 被引用 299 次
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang 等ICLR 2026 · 被引用 112 次
- Autoregressive Speech Synthesis without Vector QuantizationLingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen 等ACL 2025 · 被引用 94 次
- CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-SpeechJaehyeon Kim, Keon Lee, Seungjun Chung, Jaewoong ChoICLR 2024 · 被引用 67 次
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng 等CVPR 2026 · 被引用 40 次
它引用的顶会 Paper7
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen 等NeurIPS 2021 · 被引用 2,126 次
- Neural Networks Fail to Learn Periodic Functions and How to Fix ItLiu Ziyin, Tilman Hartwig, Masahito UedaNeurIPS 2020 · 被引用 249 次
- Chunked Autoregressive GAN for Conditional Waveform SynthesisMax Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman 等ICLR 2022 · 被引用 91 次
相关 Paper
- Toward Complex-Valued Neural Networks for Waveform GenerationHyung-Seok Oh, Deok-Hyeon Cho, Seung-Bin Kim, Seong-Whan LeeICLR 2026
- DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveAndong Li, Tong Lei, Lingling Dai, Kai Li 等AAAI 2026
- Avocodo: Generative Adversarial Network for Artifact-Free VocoderTaejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang 等AAAI 2023 · 被引用 44 次
- Representing Sounds as Neural Amplitude Fields: A Benchmark of Coordinate-MLPs and a Fourier Kolmogorov-Arnold FrameworkLinfei Li, Lin Zhang, Zhong Wang, Fengyi Zhang 等AAAI 2025 · 被引用 3 次
- Beyond Regular Grids: Fourier-Based Neural Operators on Arbitrary DomainsLevi E. Lingsch, Mike Yan Michelis, Emmanuel de Bézenac, Sirani M. Perera 等ICML 2024 · 被引用 24 次
