Controllable Music Loops Generation with MIDI and Text via Multi-Stage Cross Attention and Instrument-Aware Reinforcement Learning
Guan-Yuan Chen, Von-Wun Soo
摘要
The burgeoning field of text-to-music generation models has shown great promise in their ability to generate high-quality music aligned with users' textual descriptions. These models effectively capture abstract/global musical features such as style and mood. However, they often inadequately produce the precise rendering of critical music loop attributes, including melody, rhythms, and instrumentation, which are essential for modern music loop production. To overcome this limitation, this paper proposed a Loops Transformer and a Multi-Stage Cross Attention mechanism that enable a cohesive integration of textual and MIDI input specifications. Additionally, a novel Instrument-Aware Reinforcement Learning technique was introduced to ensure the correct adoption of instrumentation. We demonstrated that the proposed model can generate music loops that simultaneously satisfy the conditions specified by both natural language texts and MIDI input, ensuring coherence between the two modalities. We also showed that our model outperformed the state-of-the-art baseline model, MusicGen, in both objective metrics (by lowering the FAD score by 1.3, indicating superior quality with lower scores, and by improving the Normalized Dynamic Time Warping Distance with given melodies by 12%) and subjective metrics (by +2.56% in OVL, +5.42% in REL, and +7.74% in Loop Consistency). These improvements highlight our model's capability to produce musically coherent loops that satisfy the complex requirements of contemporary music production, representing a notable advancement in the field. Generated music loop samples can be explored at: https://loopstransformer.netlify.app/.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music GenerationZihao Zhang, Xingjiao Wu, Junjie Xu, Tianlong Ma 等ACM MM 2025
- MIDILM: A Dual-Path Model for Controllable Text-to-MIDI GenerationShuyu Li, Dooho Choi, Yunsick SungAAAI 2026
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music ModelingZixun Guo, Jaeyong Kang, Dorien HerremansAAAI 2023 · 被引用 27 次
- Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion ModelsYi Yang, Haowen Li, Tianxiang Li, Boyu Cao 等AAAI 2026 · 被引用 1 次
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent 等ICML 2024 · 被引用 41 次
