Masked Audio Generation using a Single Non-Autoregressive Transformer
Alon Ziv, Itai Gat, Gaël Le Lan, Tal Remez, Felix Kreuk, Jade Copet, Alexandre Défossez, Gabriel Synnaeve, Yossi Adi
Abstract
We introduce MAGNET, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAG-NET is comprised of a single-stage, non-autoregressive transformer. During training, we predict spans of masked tokens obtained from a masking scheduler, while during inference we gradually construct the output sequence using several decoding steps. To further enhance the quality of the generated audio, we introduce a novel rescoring method in which, we leverage an external pretrained model to rescore and rank predictions from MAGNET, which will be then used for later decoding steps. Lastly, we explore a hybrid version of MAG-NET, in which we fuse between autoregressive and non-autoregressive models to generate the first few seconds in an autoregressive manner while the rest of the sequence is being decoded in parallel. We demonstrate the efficiency of MAGNET for the task of text-to-music and text-to-audio generation and conduct an extensive empirical evaluation, considering both objective metrics and human studies. The proposed approach is comparable to the evaluated baselines, while being significantly faster (x7 faster than the autoregressive baseline). Through ablation studies and analysis, we shed light on the importance of each of the components comprising MAGNET, together with pointing to the trade-offs between autoregressive and non-autoregressive modeling, considering latency, throughput, and generation quality. Samples are available on our demo page https://pages.cs.huji.ac.il/adiyoss-lab/MAGNeT * Work was done as part of Alon's internship at FAIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2aef3946-b870-45a4-aeff-02bf04d93d30Cited by top-tier papers24
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- Behavior Generation with Latent ActionsSeungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim et al.ICML 2024 · 154 citations
- LaViDa: A Large Diffusion Language Model for Multimodal UnderstandingShufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul et al.NeurIPS 2025 · 89 citations
- D-Flow: Differentiating through Flows for Controlled GenerationHeli Ben-Hamu, Omri Puny, Itai Gat, Brian Karrer et al.ICML 2024 · 82 citations
- Tell What You Hear From What You See - Video to Audio Generation Through TextXiulong Liu, Kun Su, Eli ShlizermanNeurIPS 2024 · 46 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
Related papers
- IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion ModelingKuan-Po Huang, Shu-Wen Yang, Huy Phan, Bo-Ru Lu et al.ICML 2025
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng et al.ICLR 2025
- From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint TrainingTianqiao Liu, Xueyi Li, Hao Wang, Haoxuan Li et al.ICLR 2026 · 6 citations
- Non-Autoregressive Neural Text-to-SpeechKainan Peng, Wei Ping, Zhao Song, Kexin ZhaoICML 2020 · 118 citations
- Scaling Transformers for End-to-End Discrete Audio TokenizationYitian Gong, Kuangwei Chen, Zhaoye Fei, Xiaogui Yang et al.ICML 2026
