RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement Learning
Nan Jiang, Sheng Jin, Zhiyao Duan, Changshui Zhang
Abstract
This paper presents a deep reinforcement learning algorithm for online accompaniment generation, with potential for real-time interactive human-machine duet improvisation. Different from offline music generation and harmonization, online music accompaniment requires the algorithm to respond to human input and generate the machine counterpart in a sequential order. We cast this as a reinforcement learning problem, where the generation agent learns a policy to generate a musical note (action) based on previously generated context (state). The key of this algorithm is the well-functioning reward model. Instead of defining it using music composition rules, we learn this model from monophonic and polyphonic training data. This model considers the compatibility of the machine-generated note with both the machine-generated context and the human-generated context. Experiments show that this algorithm is able to respond to the human part and generate a melodic, harmonic and diverse machine part. Subjective evaluations on preferences show that the proposed algorithm generates music pieces of higher quality than the baseline method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2024c96-c5ca-4e6e-b731-6c10497f8cfaCited by top-tier papers9
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent et al.ICML 2024 · 41 citations
- When Counterpoint Meets Chinese Folk MelodiesNan Jiang, Sheng Jin, Zhiyao Duan, Changshui ZhangNeurIPS 2020 · 13 citations
- SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasZihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang et al.ACM MM 2022 · 12 citations
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music InteractionYusong Wu, Stephen Brade, Teng Ma, Tia-Jane Fowler et al.ICLR 2026 · 4 citations
- Adaptive Accompaniment with ReaLchordsYusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts et al.ICML 2024 · 4 citations
Related papers
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance AccompanimentLi Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin et al.ICLR 2024 · 54 citations
- Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue ModelsYifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen et al.ACL 2026 · 7 citations
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 54 citations
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-TrainingRan Xu, Tianci Liu, Zihan Dong, Tony Yu et al.ICML 2026
