FIGARO: Controllable Music Generation using Learned and Expert Features
Dimitri von Rütte, Luca Biggio, Yannic Kilcher, Thomas Hofmann
Abstract
Recent symbolic music generative models have achieved significant improvements in the quality of the generated samples. Nevertheless, it remains hard for users to control the output in such a way that it matches their expectation. To address this limitation, high-level, human-interpretable conditioning is essential. In this work, we release FIGARO, a Transformer-based conditional model trained to generate symbolic music based on a sequence of high-level control codes. To this end, we propose description-to-sequence learning, which consists of automatically extracting fine-grained, human-interpretable features (the description) and training a sequence-to-sequence model to reconstruct the original sequence given only the description as input. FIGARO achieves state-of-the-art performance in multi-track symbolic music generation both in terms of style transfer and sample quality. We show that performance can be further improved by combining human-interpretable with learned features. Our extensive experimental evaluation shows that FIGARO is able to generate samples that closely adhere to the content of the input descriptions, even when they deviate significantly from the training distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a8f373e-5cd3-4f26-9f38-e3ca1d5e552fCited by top-tier papers4
- Text2midi: Generating Symbolic Music from CaptionsKeshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri et al.AAAI 2025 · 21 citations
- MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music CompositionPhilippe Pasquier, Jeff Ens, Nathan Fradet, Paul Triana et al.AAAI 2025 · 14 citations
- Byte Pair Encoding for Symbolic MusicNathan Fradet, Nicolas Gutowski, Fabien Chhel, Jean-Pierre BriotEMNLP 2023 · 8 citations
- CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical ControlsLi Chai, Donglin WangAAAI 2025 · 1 citation
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Deep Learning For Symbolic MathematicsGuillaume Lample, François ChartonICLR 2020 · 477 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
Related papers
- Encoding Musical Style with Transformer AutoencodersKristy Choi, Curtis Hawthorne, Ian Simon, Monica Dinculescu et al.ICML 2020 · 102 citations
- Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic MusicHongju Su, Ke Li, Lan Yang, Honggang Zhang et al.ACL 2026
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music GenerationTingyu Zhu, Haoyu Liu, Ziyu Wang, Zhimin Jiang et al.ICML 2025
- Bridging Piano Transcription and Rendering via Disentangled Score Content and StyleWei Zeng, Junchuan Zhao, Ye WangICLR 2026
