Catch-A-Waveform: Learning to Generate Audio from a Single Short Example
Gal Greshler, Tamar Rott Shaham, Tomer Michaeli
Abstract
Models for audio generation are typically trained on hours of recordings. Here, we illustrate that capturing the essence of an audio source is typically possible from as little as a few tens of seconds from a single training signal. Specifically, we present a GAN-based generative model that can be trained on one short audio signal from any domain (e.g. speech, music, etc.) and does not require pre-training or any other form of external supervision. Once trained, our model can generate random samples of arbitrary duration that maintain semantic similarity to the training waveform, yet exhibit new compositions of its audio primitives. This enables a long line of interesting applications, including generating new jazz improvisations or new a-cappella rap variants based on a single short example, producing coherent modifications to famous songs (e.g. adding a new verse to a Beatles song based solely on the original recording), filling-in of missing parts (inpainting), extending the bandwidth of a speech signal (super-resolution), and enhancing old recordings without access to any clean training example. We show that in all cases, no more than 20 seconds of training audio commonly suffice for our model to achieve state-of-the-art results. This is despite its complete lack of prior knowledge about the nature of audio signals in general.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95263aea-a638-4a80-b937-17fceb0c79daCited by top-tier papers4
- SinDDM: A Single Image Denoising Diffusion ModelVladimir Kulikov, Shahar Yadin, Matan Kleiner, Tomer MichaeliICML 2023 · 113 citations
- Sin3DM: Learning a Diffusion Model from a Single 3D Textured ShapeRundi Wu, Ruoshi Liu, Carl Vondrick, Changxi ZhengICLR 2024 · 34 citations
- Dilated convolution with learnable spacingsIsmail Khalfaoui Hassani, Thomas Pellegrini, Timothée MasquelierICLR 2023 · 18 citations
- Moûsai: Efficient Text-to-Music Diffusion ModelsFlavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard SchölkopfACL 2024
Builds on9
- SinGAN: Learning a Generative Model From a Single Natural ImageTamar Rott Shaham, Tali Dekel, Tomer MichaeliICCV 2019 · 933 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- DDSP: Differentiable Digital Signal ProcessingJesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam RobertsICLR 2020 · 467 citations
- High Fidelity Speech Synthesis with Adversarial NetworksMikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark et al.ICLR 2020 · 263 citations
- WaveFlow: A Compact Flow-based Model for Raw AudioWei Ping, Kainan Peng, Kexin Zhao, Zhao SongICML 2020 · 132 citations
Related papers
- Chunked Autoregressive GAN for Conditional Waveform SynthesisMax Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman et al.ICLR 2022 · 91 citations
- Token-Based Audio Inpainting via Discrete DiffusionTali Dror, Iftach Shoham, Moshe Buchris, Oren Gal et al.ICLR 2026 · 3 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Listening to Sounds of Silence for Speech DenoisingRuilin Xu, Rundi Wu, Yuko Ishiwaka, Carl Vondrick et al.NeurIPS 2020 · 40 citations
