Language-Guided Music Recommendation for Video via Prompt Analogies
Daniel McKee, Justin Salamon, Josef Sivic, Bryan C. Russell
摘要
We propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting is that existing music video datasets provide the needed (video, music) training pairs, but lack text descriptions of the music. This work addresses this challenge with the following three contributions. First, we propose a textsynthesis approach that relies on an analogy-based prompting procedure to generate natural language music descriptions from a large-scale language model (BLOOM-176B) given pre-trained music tagger outputs and a small number of human text descriptions. Second, we use these synthesized music descriptions to train a new trimodal model, which fuses text and video input representations to query music samples. For training, we introduce a text dropout regularization mechanism which we show is critical to model performance. Our model design allows for the retrieved music audio to agree with the two input modalities by matching visual style depicted in the video and musical genre, mood, or instrumentation described in the natural language query. Third, to evaluate our approach, we collect a testing dataset for our problem by annotating a subset of 4k clips from the YT8M-MusicVideo dataset with natural language music descriptions which we make publicly available. We show that our approach can match or exceed the performance of prior methods on video-to-music retrieval while significantly improving retrieval accuracy when using text guidance. * Work done as an intern with Adobe Research Upbeat pop Language-Guided Music Retrieval A somber folk ballad featuring female vocalist, guitar, and tambourine Video+Language Query Retrieved Music Upbeat pop ViML Pop featuring a female vocalist with energetic synth melodies Lighthearted pop song with a male vocalist backed by guitar, drums ViML ViML Folk music with guitar Descriptions of the retrieved music tracks provided by the authors for illustration only. User input includes video and a text description of target music.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLMChangli Tang, Qinfan Xiao, Ke Mei, Tianyi Wang 等ICLR 2026 · 被引用 9 次
- Music Grounding by Short VideoZijie Xin, Minquan Wang, Jingyu Liu, Quan Chen 等ICCV 2025 · 被引用 2 次
- HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal SynchronizationZitang Zhou, Ke Mei, Yu Lu, Tianyi Wang 等CVPR 2025
- FilmComposer: LLM-Driven Music Production for Silent Film ClipsZhifeng Xie, Qile He, Youjia Zhu, Qiwei He 等CVPR 2025
- MuseChat: A Conversational Music Recommendation System for VideosZhikang Dong, Xiulong Liu, Bin Chen, Pawel Polak 等CVPR 2024
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
相关 Paper
- VMChill: A Dataset for Fine-Grained Visual-Musical SynergyXiaowei Chi, Zeyue Tian, Jialiang Chen, Wei XueAAAI 2026
- VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term ModelingZeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan 等CVPR 2025
- Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMZinuo Li, Xian Zhang, Yongxin Guo, Mohammed Bennamoun 等NeurIPS 2025 · 被引用 9 次
- Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM AnimatorHanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu 等NeurIPS 2023 · 被引用 111 次
- Multi-Modal Inductive Framework for Text-Video RetrievalQian Li, Yucheng Zhou, Cheng Ji, Feihong Lu 等ACM MM 2024 · 被引用 8 次
