From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion
Robin San Roman, Yossi Adi, Antoine Deleforge, Romain Serizel, Gabriel Synnaeve, Alexandre Défossez
Abstract
Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms conditioned on highly compressed representations. Although such methods produce impressive results, they are prone to generate audible artifacts when the conditioning is flawed or imperfect. An alternative modeling approach is to use diffusion models. However, these have mainly been used as speech vocoders (i.e., conditioned on mel-spectrograms) or generating relatively low sampling rate signals. In this work, we propose a high-fidelity multi-band diffusion-based framework that generates any type of audio modality (e.g., speech, music, environmental sounds) from low-bitrate discrete representations. At equal bit rate, the proposed approach outperforms state-of-the-art generative techniques in terms of perceptual quality. Training and evaluation code are available on the facebookresearch/audiocraft github project. Samples are available on the following link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab9a744f-3896-4cc0-9040-d820db752a16Cited by top-tier papers10
- Latent Fourier TransformMason Wang, Cheng-Zhi Anna HuangICLR 2026 · 58 citations
- Audio Super-Resolution with Latent Bridge ModelsChang Li, Zehua Chen, Liyuan Wang, Jun ZhuNeurIPS 2025 · 18 citations
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio GenerationZengwei Yao, Wei Kang, Han Zhu, Liyong Guo et al.ICLR 2026 · 5 citations
- StreamFlow: Streaming Audio Generation from Discrete Tokens via Streaming Flow MatchingHa-Yeong Choi, Sang-Hoon LeeNeurIPS 2025 · 2 citations
- FlowDec: A flow-based full-band general audio codec with high perceptual qualitySimon Welker, Matthew Le, Ricky T. Q. Chen, Wei-Ning Hsu et al.ICLR 2025
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- WaveGrad: Estimating Gradients for Waveform GenerationNanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss et al.ICLR 2021 · 44 citations
- MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video GenerationMingzhen Sun, Weining Wang, Yanyuan Qiao, Jiahui Sun et al.ACM MM 2024 · 4 citations
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
