RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement Learning
Nan Jiang, Sheng Jin, Zhiyao Duan, Changshui Zhang
摘要
This paper presents a deep reinforcement learning algorithm for online accompaniment generation, with potential for real-time interactive human-machine duet improvisation. Different from offline music generation and harmonization, online music accompaniment requires the algorithm to respond to human input and generate the machine counterpart in a sequential order. We cast this as a reinforcement learning problem, where the generation agent learns a policy to generate a musical note (action) based on previously generated context (state). The key of this algorithm is the well-functioning reward model. Instead of defining it using music composition rules, we learn this model from monophonic and polyphonic training data. This model considers the compatibility of the machine-generated note with both the machine-generated context and the human-generated context. Experiments show that this algorithm is able to respond to the human part and generate a melodic, harmonic and diverse machine part. Subjective evaluations on preferences show that the proposed algorithm generates music pieces of higher quality than the baseline method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent 等ICML 2024 · 被引用 41 次
- When Counterpoint Meets Chinese Folk MelodiesNan Jiang, Sheng Jin, Zhiyao Duan, Changshui ZhangNeurIPS 2020 · 被引用 13 次
- SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasZihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang 等ACM MM 2022 · 被引用 12 次
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music InteractionYusong Wu, Stephen Brade, Teng Ma, Tia-Jane Fowler 等ICLR 2026 · 被引用 4 次
- Adaptive Accompaniment with ReaLchordsYusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts 等ICML 2024 · 被引用 4 次
相关 Paper
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson 等EMNLP 2020 · 被引用 9 次
- Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance AccompanimentLi Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin 等ICLR 2024 · 被引用 54 次
- Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue ModelsYifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen 等ACL 2026 · 被引用 7 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-TrainingRan Xu, Tianci Liu, Zihan Dong, Tony Yu 等ICML 2026
