Simple and Controllable Music Generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, Alexandre Défossez
摘要
We tackle the task of conditional music generation. We introduce MUSICGEN, a single Language Model (LM) that operates over several streams of compressed discrete music representation, i.e., tokens. Unlike prior work, MUSICGEN is comprised of a single-stage transformer LM together with efficient token interleaving patterns, which eliminates the need for cascading several models, e.g., hierarchically or upsampling. Following this approach, we demonstrate how MUSICGEN can generate high-quality samples, both mono and stereo, while being conditioned on textual description or melodic features, allowing better controls over the generated output. We conduct extensive empirical evaluation, considering both automatic and human studies, showing the proposed approach is superior to the evaluated baselines on a standard text-to-music benchmark. Through ablation studies, we shed light over the importance of each of the components comprising MUSICGEN. Music samples, code, and models are available at github.com/facebookresearch/audiocraft.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper153
- Discrete Flow MatchingItai Gat, Tal Remez, Neta Shaul, Felix Kreuk 等NeurIPS 2024 · 被引用 363 次
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley 等ICML 2024 · 被引用 220 次
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin 等ICML 2024 · 被引用 124 次
- Proactive Detection of Voice Cloning with Localized WatermarkingRobin San Roman, Pierre Fernandez, Hady Elsahar, Alexandre Défossez 等ICML 2024 · 被引用 119 次
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang 等ICLR 2026 · 被引用 112 次
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
相关 Paper
- MusicFlow: Cascaded Flow Matching for Text Guided Music GenerationK. R. Prajwal, Bowen Shi, Matthew Le, Apoorv Vyas 等ICML 2024 · 被引用 18 次
- Continuous Audio Language ModelsSimon Rouard, Manu Orsini, Axel Roebel, Neil Zeghidour 等ICLR 2026 · 被引用 13 次
- Efficient Neural Music GenerationMax W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin 等NeurIPS 2023 · 被引用 95 次
- AudioGen: Textually Guided Audio GenerationFelix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer 等ICLR 2023 · 被引用 54 次
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token PredictionShu-Wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang 等ICML 2025
