Compose with Me: Collaborative Music Inpainter for Symbolic Music Infilling
Zhejing Hu, Yan Liu, Gong Chen, Bruce X. B. Yu
Abstract
The field of music generation has seen a surge of interest from both academia and industry, with innovative platforms such as Suno, Udio, and SkyMusic earning widespread recognition. However, the challenge of music infilling—modifying specific music segments without reconstructing the entire piece—remains a significant hurdle for both audio-based and symbolic-based models, limiting their adaptability and practicality. In this paper, we address symbolic music infilling by introducing the Collaborative Music Inpainter (CMI), an advanced human-in-the-loop (HITL) model for music infilling. The CMI features the Joint Embedding Predictive Autoregressive Generative Architecture (JEP-AGA), which learns the high-level predictive representations of the masked part that needs to be infilled during the autoregressive generative process, akin to how humans perceive and interpret music. The newly developed Dynamic Interaction Learner (DIL) achieves HITL by iteratively refining the infilled output based on user interactions alone, significantly reducing the interaction cost without requiring further input. Experimental results confirm CMI’s superior performance in music infilling, demonstrating its efficiency in producing high-quality music.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c5197a4-2621-4d8a-9af1-48abff49a619Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu et al.NeurIPS 2022 · 104 citations
- Efficient Neural Music GenerationMax W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin et al.NeurIPS 2023 · 95 citations
- MARTA: Leveraging Human Rationales for Explainable Text ClassificationInes Arous, Ljiljana Dolamic, Jie Yang, Akansha Bhardwaj et al.AAAI 2021 · 47 citations
Related papers
- S²MILE: Semantic-and-Structure-Aware Music-Driven Lyric GenerationMu You, Fang Zhang, Shuai Zhang, Linli XuAAAI 2025
- CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical ControlsLi Chai, Donglin WangAAAI 2025 · 1 citation
- MIDILM: A Dual-Path Model for Controllable Text-to-MIDI GenerationShuyu Li, Dooho Choi, Yunsick SungAAAI 2026
- MusicFlow: Cascaded Flow Matching for Text Guided Music GenerationK. R. Prajwal, Bowen Shi, Matthew Le, Apoorv Vyas et al.ICML 2024 · 18 citations
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent et al.ICML 2024 · 41 citations
