Moûsai: Efficient Text-to-Music Diffusion Models
Flavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard Schölkopf
摘要
Recent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -music. Music, much like text, can convey emotions, stories, and ideas, and has its own unique structure and syntax. In our work, we bridge text and music via a textto-music generation model that is highly efficient, expressive, and can handle long-term structure. Specifically, we develop Moûsai, a cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions. Moreover, our model features high efficiency, which enables real-time inference on a single consumer GPU with a reasonable speed. Through experiments and property analyses, we show our model's competence over a variety of criteria compared with existing music generation models. Lastly, to promote the opensource culture, we provide a collection of opensource libraries with the hope of facilitating future work in the field. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Latent Fourier TransformMason Wang, Cheng-Zhi Anna HuangICLR 2026 · 被引用 58 次
- Audio Super-Resolution with Latent Bridge ModelsChang Li, Zehua Chen, Liyuan Wang, Jun ZhuNeurIPS 2025 · 被引用 18 次
- Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion TransformersWongi Jeong, Kyungryeol Lee, Hoigi Seo, Se Young ChunCVPR 2026 · 被引用 10 次
- MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source ExtractionYunkee Chae, Kyogu LeeNeurIPS 2025 · 被引用 4 次
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 被引用 276 次
相关 Paper
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley 等ICML 2024 · 被引用 220 次
- Efficient Neural Music GenerationMax W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin 等NeurIPS 2023 · 被引用 95 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- Text2midi: Generating Symbolic Music from CaptionsKeshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri 等AAAI 2025 · 被引用 21 次
- Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion ModelsZiyu Wang, Lejun Min, Gus XiaICLR 2024 · 被引用 32 次
