SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure Bias
Zihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang, Qihao Liang, Pengfei Yu, Yongsheng Feng, Wenbo Liu, Yikai Wang, Yuntai Bao, Yiheng Yang
摘要
Real-time music accompaniment generation has a wide range of applications in the music industry, such as music education and live performances. However, automatic real-time music accompaniment generation is still understudied and often faces a trade-off between logical latency and exposure bias. In this paper, we propose SongDriver, a real-time music accompaniment generation system without logical latency nor exposure bias. Specifically, SongDriver divides one accompaniment generation task into two phases: 1) The arrangement phase, where a Transformer model first arranges chords for input melodies in real-time, and caches the chords for the next phase instead of playing them out. 2) The prediction phase, where a CRF model generates playable multi-track accompaniments for the coming melodies based on previously cached chords. With this two-phase strategy, SongDriver directly generates the accompaniment for the upcoming melody, achieving zero logical latency. Furthermore, when predicting chords for a timestep, SongDriver refers to the cached chords from the first phase rather than its previous predictions, which avoids the exposure bias problem. Since the input length is often constrained under real-time conditions, another potential problem is the loss of long-term sequential information. To make up for this disadvantage, we extract four musical features from a long-term music piece before the current time step as global information. In the experiment, we train SongDriver on some open-source datasets and an original àiMusic Dataset built from Chinese-style modern pop music sheets. The results show that SongDriver outperforms existing SOTA (state-of-the-art) models on both objective and subjective metrics, meanwhile significantly reducing the physical latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Amuse: Human-AI Collaborative Songwriting with Multimodal InspirationsYewon Kim, Sung-Ju Lee, Chris DonahueCHI 2025 · 被引用 35 次
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music InteractionYusong Wu, Stephen Brade, Teng Ma, Tia-Jane Fowler 等ICLR 2026 · 被引用 4 次
- BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal StepsLekai Qian, Haoyu Gu, Jingwei Zhao, Ziyu WangICML 2026 · 被引用 3 次
- A Design Space for Live Music AgentsYewon Kim, Stephen Brade, Alexander Wang, David Zhou 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper4
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 被引用 242 次
- PopMAG: Pop Music Accompaniment GenerationYi Ren, Jinzheng He, Xu Tan, Tao Qin 等ACM MM 2020 · 被引用 91 次
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang 等ACL 2020 · 被引用 81 次
- RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement LearningNan Jiang, Sheng Jin, Zhiyao Duan, Changshui ZhangAAAI 2020 · 被引用 56 次
相关 Paper
- Adaptive Accompaniment with ReaLchordsYusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts 等ICML 2024 · 被引用 4 次
- SongCreator: Lyrics-based Universal Song GenerationShun Lei, Yixuan Zhou, Boshi Tang, Max W. Y. Lam 等NeurIPS 2024 · 被引用 33 次
- Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and AccompanimentZhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang 等ACL 2024 · 被引用 3 次
- SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-TrainingJiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang 等AAAI 2025 · 被引用 2 次
- SongMASS: Automatic Song Writing with Pre-training and Alignment ConstraintZhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren 等AAAI 2021 · 被引用 84 次
