A Spectral Energy Distance for Parallel Speech Synthesis
Alexey A. Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek, Nal Kalchbrenner
摘要
Speech synthesis is an important practical generative modeling problem that has seen great progress over the last few years, with likelihood-based autoregressive neural models now outperforming traditional concatenative systems. A downside of such autoregressive models is that they require executing tens of thousands of sequential operations per second of generated audio, making them ill-suited for deployment on specialized deep learning hardware. Here, we propose a new learning method that allows us to train highly parallel models of speech, without requiring access to an analytical likelihood function. Our approach is based on a generalized energy distance between the distributions of the generated and real audio. This spectral energy distance is a proper scoring rule with respect to the distribution over magnitude-spectrograms of the generated waveform audio and offers statistical consistency guarantees. The distance can be calculated from minibatches without bias, and does not involve adversarial learning, yielding a stable and consistent method for training implicit generative models. Empirically, we achieve state-of-the-art generation quality among implicit generative models, as judged by the recently-proposed cFDSD metric. When combining our method with adversarial techniques, we also improve upon the recently-proposed GAN-TTS model in terms of Mean Opinion Score as judged by trained human evaluators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesisHubert SiuzdakICLR 2024 · 被引用 229 次
- Proactive Detection of Voice Cloning with Localized WatermarkingRobin San Roman, Pierre Fernandez, Hady Elsahar, Alexandre Défossez 等ICML 2024 · 被引用 119 次
- Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale CorpusRongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu 等ACM MM 2021 · 被引用 75 次
- Non-adversarial training of Neural SDEs with signature kernel scoresZacharia Issa, Blanka Horvath, Maud Lemercier, Cristopher SalviNeurIPS 2023 · 被引用 56 次
它引用的顶会 Paper1
相关 Paper
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceZhengrui Ma, Yang Feng, Chenze Shao, Fandong Meng 等NeurIPS 2025 · 被引用 5 次
- Chunked Autoregressive GAN for Conditional Waveform SynthesisMax Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman 等ICLR 2022 · 被引用 91 次
- TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech SystemsChristoph Minixhofer, Ondrej Klejch, Peter BellICLR 2026 · 被引用 16 次
- A Characteristic Function Approach to Deep Implicit Generative ModelingAbdul Fatir Ansari, Jonathan Scarlett, Harold SohCVPR 2020
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim 等AAAI 2021 · 被引用 60 次
