Diff-BGM: A Diffusion Model for Video Background Music Generation
Sizhe Li, Yiming Qin, Minghang Zheng, Xin Jin, Yang Liu
Abstract
When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the difficulties in flexibly controlling the music generation process and sequentially aligning the video and music. In this work, we first propose a high-quality music-video dataset BGM909 with detailed annotation and shot detection to provide multimodal information about the video and music. We then present evaluation metrics to assess music quality, including music diversity and alignment between music and video with retrieval precision metrics. Finally, we propose the Diff-BGM framework to automatically generate the background music for a given video, which uses different signals to control different aspects of the music during the generation process, i.e., uses dynamic video features to control music rhythm and semantic features to control the melody and atmosphere. We propose to align the video and music sequentially by introducing a segment-aware crossattention layer. Experiments verify the effectiveness of our proposed method. The code and models are available at https://github.com/sizhelee/Diff-BGM .
"a large plume of smoke coming from an open container" Captions Videos Videos with BGM k Melody k Rhythm 𝑡𝑡 𝑁𝑁 𝑡𝑡 𝑖𝑖 𝑡𝑡 0
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d24f490f-68b7-46bd-acb4-fd3f16e2b448Cited by top-tier papers15
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan et al.ICLR 2026 · 38 citations
- GVMGen: A General Video-to-Music Generation Model with Hierarchical AttentionsHeda Zuo, Weitao You, Junxian Wu, Shihong Ren et al.AAAI 2025 · 15 citations
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music GenerationXinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan et al.AAAI 2026 · 4 citations
- Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation LearningJiayun Hu, Yueyi He, Tianyi Liang, Changbo Wang et al.ACM MM 2025 · 2 citations
- Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation ModelsChristian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito et al.CVPR 2026 · 2 citations
Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
Related papers
- Video Background Music Generation with Controllable Music TransformerShangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang et al.ACM MM 2021 · 87 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
- Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music GenerationShulei Ji, Zihao Wang, Jiaxing Yu, Xiangyuan Yang et al.AAAI 2026
- VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term ModelingZeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan et al.CVPR 2025
- Audio-Sync Video Generation with Multi-Stream Temporal ControlShuchen Weng, Haojie Zheng, Zheng Chang, Si Li et al.NeurIPS 2025 · 14 citations
