BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, Sungroh Yoon
摘要
Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across various recording environments. In this work, we present BigVGAN, a universal vocoder that generalizes well for various out-of-distribution scenarios without fine-tuning. We introduce periodic activation function and anti-aliased representation into the GAN generator, which brings the desired inductive bias for audio synthesis and significantly improves audio quality. In addition, we train our GAN vocoder at the largest scale up to 112M parameters, which is unprecedented in the literature. We identify and address the failure modes in large-scale GAN training for audio, while maintaining high-fidelity output without over-regularization. Our BigVGAN, trained only on clean speech (LibriTTS), achieves the state-of-the-art performance for various zero-shot (out-of-distribution) conditions, including unseen speakers, languages, recording environments, singing voices, music, and instrumental audio. 1 We release our code and model at: https://github.com/NVIDIA/BigVGAN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsHeng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright 等AAAI 2024 · 被引用 84 次
- Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow MatchingYongqi Wang, Wenxiang Guo, Rongjie Huang, Jiawei Huang 等NeurIPS 2024 · 被引用 73 次
- CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-SpeechJaehyeon Kim, Keon Lee, Seungjun Chung, Jaewoong ChoICLR 2024 · 被引用 67 次
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou 等AAAI 2026 · 被引用 63 次
- StyleSinger: Style Transfer for Out-of-Domain Singing Voice SynthesisYu Zhang, Rongjie Huang, Ruiqi Li, Jinzheng He 等AAAI 2024 · 被引用 44 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
相关 Paper
- DFlow: A Generative Model Combining Denoising AutoEncoder and Normalizing Flow for High Fidelity Waveform GenerationChenfeng Miao, Qingying Zhu, Minchuan Chen, Wei Hu 等ICML 2024 · 被引用 2 次
- SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationRongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren 等ACM MM 2022 · 被引用 46 次
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri 等ACM MM 2022 · 被引用 15 次
- DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveAndong Li, Tong Lei, Lingling Dai, Kai Li 等AAAI 2026
- Unsupervised Image-to-Image Translation with Generative PriorShuai Yang, Liming Jiang, Ziwei Liu, Chen Change LoyCVPR 2022 · 被引用 51 次
