It's Time for Artistic Correspondence in Music and Video
Dídac Surís, Carl Vondrick, Bryan C. Russell, Justin Salamon
摘要
We present an approach for recommending a music track for a given video, and vice versa, based on both their temporal alignment and their correspondence at an artistic level. We propose a self-supervised approach that learns this correspondence directly from data, without any need of human annotations. In order to capture the high-level concepts that are required to solve the task, we propose modeling the long-term temporal context of both the video and the music signals, using Transformer networks for each modality. Experiments show that this approach strongly outperforms alternatives that do not exploit the temporal context. The combination of our contributions improve retrieval accuracy up to 10× over prior state of the art. This strong improvement allows us to introduce a wide range of analyses and applications. For instance, we can condition music retrieval based on visually defined attributes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao 等ICCV 2023 · 被引用 51 次
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 等AAAI 2024 · 被引用 29 次
- GVMGen: A General Video-to-Music Generation Model with Hierarchical AttentionsHeda Zuo, Weitao You, Junxian Wu, Shihong Ren 等AAAI 2025 · 被引用 15 次
- Long-range Multimodal Pretraining for Movie UnderstandingDawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon 等ICCV 2023 · 被引用 15 次
- Music Grounding by Short VideoZijie Xin, Minquan Wang, Jingyu Liu, Quan Chen 等ICCV 2025 · 被引用 2 次
相关 Paper
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 被引用 149 次
- Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music GenerationWeitao You, Heda Zuo, Junxian Wu, Dengming Zhang 等ACM MM 2025
- Aligning Moments in Time Using Video QueriesYogesh Kumar, Uday Agarwal, Manish Gupta, Anand MishraICCV 2025 · 被引用 2 次
- Audio-Visual Contrastive Learning with Temporal Self-SupervisionSimon Jenni, Alexander Black, John P. CollomosseAAAI 2023 · 被引用 25 次
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
