Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation
Shulei Ji, Zihao Wang, Jiaxing Yu, Xiangyuan Yang, Shuyu Li, Songruoyao Wu, Kejun Zhang
摘要
Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2) effectively integrating various visual features to condition music generation remains non-trivial. To address these issues, we propose Diff-V2M, a general V2M framework based on a hierarchical conditional diffusion model, comprising two core components: visual feature extraction and conditional music generation. For rhythm modeling, we begin by evaluating several rhythmic representations, including lowresolution mel-spectrograms, tempograms, and onset detection functions (ODF), and devise a rhythmic predictor to infer them directly from videos. To ensure contextual and affective coherence, we also extract semantic and emotional features. All features are incorporated into the generator via a hierarchical cross-attention mechanism, where emotional features shape the affective tone via the first layer, while semantic and rhythmic features are fused in the second cross-attention layer. To enhance feature integration, we introduce timestepaware fusion strategies, including feature-wise linear modulation (FiLM) and weighted fusion, allowing the model to adaptively balance semantic and rhythmic cues throughout the diffusion process. Extensive experiments identify lowresolution ODF as a more effective signal for modeling musical rhythm and demonstrate that Diff-V2M outperforms existing models on both in-domain and out-of-domain datasets, achieving state-of-the-art performance in terms of objective metrics and subjective comparisons. Demo and code are available at https://Tayjsl97.github.io/Diff-V2M-Demo/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi 等ICML 2020 · 被引用 553 次
相关 Paper
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music GenerationXinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan 等AAAI 2026 · 被引用 4 次
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 等AAAI 2024 · 被引用 29 次
- Diff-BGM: A Diffusion Model for Video Background Music GenerationSizhe Li, Yiming Qin, Minghang Zheng, Xin Jin 等CVPR 2024
- Long-Term Rhythmic Video SoundtrackerJiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun 等ICML 2023 · 被引用 24 次
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationMoayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 等ICCV 2025 · 被引用 3 次
