Video Background Music Generation with Controllable Music Transformer
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, Shuicheng Yan
Abstract
In this work, we address the task of video background music generation. Some previous works achieve effective music generation but are unable to generate melodious music tailored to a particular video, and none of them considers the video-music rhythmic consistency. To generate the background music that matches the given video, we first establish the rhythmic relations between video and background music. In particular, we connect timing, motion speed, and motion saliency from video with beat, simu-note density, and simu-note strength from music, respectively. We then propose CMT, a Controllable Music Transformer that enables local control of the aforementioned rhythmic features and global control of the music genre and instruments. Objective and subjective evaluations show that the generated background music has achieved satisfactory compatibility with the input videos, and at the same time, impressive music quality. Code and models are available at https://github.com/wzk1015/video-bgm-generation.
• Applied computing → Sound and music computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 610dcfc8-a352-4856-a117-c8f8286f0e61Cited by top-tier papers26
- V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsHeng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright et al.AAAI 2024 · 84 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan et al.ICLR 2026 · 38 citations
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin et al.AAAI 2024 · 29 citations
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 27 citations
Builds on6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 242 citations
- PopMAG: Pop Music Accompaniment GenerationYi Ren, Jinzheng He, Xu Tan, Tao Qin et al.ACM MM 2020 · 91 citations
Related papers
- Diff-BGM: A Diffusion Model for Video Background Music GenerationSizhe Li, Yiming Qin, Minghang Zheng, Xin Jin et al.CVPR 2024
- Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music GenerationWeitao You, Heda Zuo, Junxian Wu, Dengming Zhang et al.ACM MM 2025
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music GenerationXinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan et al.AAAI 2026 · 4 citations
- Controllable Video-to-Music Generation with Multiple Time-Varying ConditionsJunxian Wu, Weitao You, Heda Zuo, Dengming Zhang et al.ACM MM 2025 · 1 citation
- GVMGen: A General Video-to-Music Generation Model with Hierarchical AttentionsHeda Zuo, Weitao You, Junxian Wu, Shihong Ren et al.AAAI 2025 · 15 citations
