It's Time for Artistic Correspondence in Music and Video
Dídac Surís, Carl Vondrick, Bryan C. Russell, Justin Salamon
Abstract
We present an approach for recommending a music track for a given video, and vice versa, based on both their temporal alignment and their correspondence at an artistic level. We propose a self-supervised approach that learns this correspondence directly from data, without any need of human annotations. In order to capture the high-level concepts that are required to solve the task, we propose modeling the long-term temporal context of both the video and the music signals, using Transformer networks for each modality. Experiments show that this approach strongly outperforms alternatives that do not exploit the temporal context. The combination of our contributions improve retrieval accuracy up to 10× over prior state of the art. This strong improvement allows us to introduce a wide range of analyses and applications. For instance, we can condition music retrieval based on visually defined attributes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin et al.AAAI 2024 · 29 citations
- GVMGen: A General Video-to-Music Generation Model with Hierarchical AttentionsHeda Zuo, Weitao You, Junxian Wu, Shihong Ren et al.AAAI 2025 · 15 citations
- Long-range Multimodal Pretraining for Movie UnderstandingDawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon et al.ICCV 2023 · 15 citations
- Music Grounding by Short VideoZijie Xin, Minquan Wang, Jingyu Liu, Quan Chen et al.ICCV 2025 · 2 citations
Related papers
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 149 citations
- Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music GenerationWeitao You, Heda Zuo, Junxian Wu, Dengming Zhang et al.ACM MM 2025
- Aligning Moments in Time Using Video QueriesYogesh Kumar, Uday Agarwal, Manish Gupta, Anand MishraICCV 2025 · 2 citations
- Audio-Visual Contrastive Learning with Temporal Self-SupervisionSimon Jenni, Alexander Black, John P. CollomosseAAAI 2023 · 25 citations
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
