VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails
Mina Huh, C. Ailie Fraser, Dingzeyu Li, Mira Dontcheva, Bryan Wang
Abstract
Music shapes the tone of videos, yet creators find it hard to find soundtracks that match their video’s mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, and understand their impact on the video. We present VidTune, a system that supports soundtrack creation by generating diverse music options from a creator’s prompt and producing contextual thumbnails for rapid review. VidTune extracts representative video subjects to ground thumbnails in context, maps each track’s valence and energy onto visual cues like color and brightness, and depicts prominent genres and instruments. Creators can refine tracks with natural language edits, which VidTune expands into new generations. In a controlled user study (N=12) and an exploratory case study (N=6), participants found VidTune helpful for efficiently reviewing and comparing music options and described the process as playful and enriching.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b85cb76e-9d88-489e-b061-b3f5419c3c7dBuilds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative ModelsRyan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry et al.CHI 2020 · 265 citations
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore et al.UIST 2023 · 179 citations
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li et al.CHI 2024 · 143 citations
- FashionQ: An AI-Driven Creativity Support Tool for Facilitating Ideation in Fashion DesignYoungseung Jeon, Seungwan Jin, Patrick C. Shih, Kyungsik HanCHI 2021 · 143 citations
Related papers
- VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term ModelingZeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan et al.CVPR 2025
- "Is Text-Based Music Search Enough to Satisfy Your Needs?" A New Way to Discover Music with ImagesJeongeun Park, Hyorim Shin, Changhoon Oh, Ha Young KimCHI 2024 · 6 citations
- MVPrompt: Building Music-Visual Prompts for AI Artists to Craft Music Video Mise-en-scèneChungHa Lee, Daeho Lee, Jin-Hyuk HongCHI 2025 · 5 citations
- Vidmento: Creating Video Stories through Context-Aware Expansion with Generative VideoCatherine Yeh, Anh Truong, Mira Dontcheva, Bryan WangCHI 2026 · 1 citation
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon et al.CHI 2026 · 2 citations
