Moûsai: Efficient Text-to-Music Diffusion Models
Flavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard Schölkopf
Abstract
Recent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -music. Music, much like text, can convey emotions, stories, and ideas, and has its own unique structure and syntax. In our work, we bridge text and music via a textto-music generation model that is highly efficient, expressive, and can handle long-term structure. Specifically, we develop Moûsai, a cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions. Moreover, our model features high efficiency, which enables real-time inference on a single consumer GPU with a reasonable speed. Through experiments and property analyses, we show our model's competence over a variety of criteria compared with existing music generation models. Lastly, to promote the opensource culture, we provide a collection of opensource libraries with the hope of facilitating future work in the field. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d801fabd-a45f-4a9b-9c7c-638c21576606Cited by top-tier papers10
- Latent Fourier TransformMason Wang, Cheng-Zhi Anna HuangICLR 2026 · 58 citations
- Audio Super-Resolution with Latent Bridge ModelsChang Li, Zehua Chen, Liyuan Wang, Jun ZhuNeurIPS 2025 · 18 citations
- Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion TransformersWongi Jeong, Kyungryeol Lee, Hoigi Seo, Se Young ChunCVPR 2026 · 10 citations
- MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source ExtractionYunkee Chae, Kyogu LeeNeurIPS 2025 · 4 citations
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang et al.CVPR 2026 · 3 citations
Builds on20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 276 citations
Related papers
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- Efficient Neural Music GenerationMax W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin et al.NeurIPS 2023 · 95 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- Text2midi: Generating Symbolic Music from CaptionsKeshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri et al.AAAI 2025 · 21 citations
- Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion ModelsZiyu Wang, Lejun Min, Gus XiaICLR 2024 · 32 citations
