Long-Term Rhythmic Video Soundtracker
Jiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun, Yu Qiao
Abstract
We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods directly generating video-conditioned waveforms suffer from limited scenarios, short lengths, and unstable generation quality. To this end, we present Long-Term Rhythmic Video Soundtracker (LORIS), a novel framework to synthesize long-term conditional waveforms. Specifically, our framework consists of a latent conditional diffusion probabilistic model to perform waveform synthesis. Furthermore, a series of context-aware conditioning encoders are proposed to take temporal information into consideration for a long-term generation. Notably, we extend our model's applicability from dances to multiple sports scenarios such as floor exercise and figure skating. To perform comprehensive evaluations, we establish a benchmark for rhythmic video soundtracks including the pre-processed dataset, improved evaluation metrics, and robust generative baselines. Extensive experiments show that our model generates long-term soundtracks with state-of-the-art musical quality and rhythmic correspondence. Codes are available at https://github.com/OpenGVLab/LORIS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe82e8cf-75c4-4ceb-9dfb-427612b0d275Cited by top-tier papers10
- MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and CorrespondenceFuming You, Minghui Fang, Li Tang, Rongjie Huang et al.NeurIPS 2024 · 8 citations
- Morph: a Motion-Free Physics Optimization Framework for Human Motion GenerationZhuo Li, Mingshuang Luo, Ruibing Hou, Xin Zhao et al.ICCV 2025 · 2 citations
- HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal SynchronizationZitang Zhou, Ke Mei, Yu Lu, Tianyi Wang et al.CVPR 2025
- Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion ModelChangchang Sun, Gaowen Liu, Charles Fleming, Yan YanCVPR 2025
- FilmComposer: LLM-Driven Music Production for Silent Film ClipsZhifeng Xie, Qile He, Youjia Zhu, Qiwei He et al.CVPR 2025
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
Related papers
- Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music GenerationShulei Ji, Zihao Wang, Jiaxing Yu, Xiangyuan Yang et al.AAAI 2026
- M2PE-Diff: Music-to-Pose Encoder for Dance Video Generation Leveraging Latent Diffusion FrameworkNokap Tony ParkACM MM 2025 · 2 citations
- MusicInfuser: Making Video Diffusion Listen and DanceSusung Hong, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. SeitzCVPR 2026 · 5 citations
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music GenerationXinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan et al.AAAI 2026 · 4 citations
- Self-supervised Dance Video Synthesis Conditioned on MusicXuanchi Ren, Haoran Li, Zijian Huang, Qifeng ChenACM MM 2020 · 68 citations
