Chunked Autoregressive GAN for Conditional Waveform Synthesis
Max Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman, Aaron C. Courville, Yoshua Bengio
摘要
Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI. These systems employ deep generative models that model the waveform via either sequential (autoregressive) or parallel (non-autoregressive) sampling. Generative adversarial networks (GANs) have become a common choice for non-autoregressive waveform synthesis. However, state-of-the-art GAN-based models produce artifacts when performing mel-spectrogram inversion. In this paper, we demonstrate that these artifacts correspond with an inability for the generator to learn accurate pitch and periodicity. We show that simple pitch and periodicity conditioning is insufficient for reducing this error relative to using autoregression. We discuss the inductive bias that autoregression provides for learning the relationship between instantaneous frequency and phase, and show that this inductive bias holds even when autoregressively sampling large chunks of the waveform during each forward pass. Relative to prior state-of-the-art GAN-based models, our proposed model, Chunked Autoregressive GAN (CARGAN) reduces pitch error by 40-60%, reduces training time by 58%, maintains a fast generation speed suitable for real-time or interactive applications, and maintains or improves subjective quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesisHubert SiuzdakICLR 2024 · 被引用 229 次
- BigVGAN: A Universal Neural Vocoder with Large-Scale TrainingSang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro 等ICLR 2023 · 被引用 46 次
- Avocodo: Generative Adversarial Network for Artifact-Free VocoderTaejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang 等AAAI 2023 · 被引用 44 次
- NANSY++: Unified Voice Synthesis with Neural Analysis and SynthesisHyeong-Seok Choi, Jinhyeok Yang, Juheon Lee, Hyeongju KimICLR 2023 · 被引用 8 次
它引用的顶会 Paper8
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- High Fidelity Speech Synthesis with Adversarial NetworksMikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark 等ICLR 2020 · 被引用 263 次
- You Only Need Adversarial Supervision for Semantic Image SynthesisEdgar Schönfeld, Vadim Sushko, Dan Zhang, Juergen Gall 等ICLR 2021 · 被引用 219 次
相关 Paper
- WaveGrad: Estimating Gradients for Waveform GenerationNanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss 等ICLR 2021 · 被引用 44 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim 等AAAI 2021 · 被引用 60 次
- Revisiting Over-Smoothness in Text to SpeechYi Ren, Xu Tan, Tao Qin, Zhou Zhao 等ACL 2022
- A Spectral Energy Distance for Parallel Speech SynthesisAlexey A. Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek 等NeurIPS 2020 · 被引用 89 次
