Soundify: Matching Sound Effects to Video
David Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi, Nikolas Martelaro
Abstract
In the art of video editing, sound helps add character to an object and immerse the viewer within a space. Through formative interviews with professional editors (N=10), we found that the task of adding sounds to video can be challenging. This paper presents Soundify, a system that assists editors in matching sounds to video. Given a video, Soundify identifies matching sounds, synchronizes the sounds to the video, and dynamically adjusts panning and volume to create spatial audio. In a human evaluation study (N=889), we show that Soundify is capable of matching sounds to video out-of-the-box for a diverse range of audio categories. In a within-subjects expert study (N=12), we demonstrate the usefulness of Soundify in helping video editors match sounds to video with lighter workload, reduced task completion time, and improved usability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1c02738-a4db-44ed-9dcb-4c401b67408fCited by top-tier papers4
- SonifyAR: Context-Aware Sound Generation in Augmented RealityXia Su, Jon E. Froehlich, Eunyee Koh, Chang XiaoUIST 2024 · 12 citations
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon et al.CHI 2026 · 2 citations
- MoSound: An Interactive Tool for Generative Sound Design in Motion GraphicsJialin Huang, Prem Seetharaman, Timothy Richard Langlois, Li-Yi Wei et al.CHI 2026 · 2 citations
- Auditorily Embodied Conversational Agents: Effects of Spatialization and Situated Audio Cues on Presence and Social PerceptionYi Fei Cheng, Jarod Bloch, Alexander Wang, Andrea Bianchi et al.CHI 2026 · 1 citation
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Audeo: Audio Generation for a Silent Performance VideoKun Su, Xiulong Liu, Eli ShlizermanNeurIPS 2020 · 78 citations
- Crosscast: Adding Visuals to Audio Travel PodcastsHaijun Xia, Jennifer Jacobs, Maneesh AgrawalaUIST 2020 · 44 citations
Related papers
- Eventfulness for Interactive Video AlignmentJiatian Sun, Longxiulin Deng, Triantafyllos Afouras, Andrew Owens et al.SIGGRAPH 2023 · 5 citations
- D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX MatchingJingyu Liu, Minquan Wang, Ye Ma, Bo Wang et al.AAAI 2025 · 4 citations
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song et al.ACM MM 2024 · 6 citations
- AutoSFX: Automatic Sound Effect Generation for VideosYujia Wang, Zhongxu Wang, Hua HuangACM MM 2024 · 2 citations
- What's Making That Sound Right Now? Video-Centric Audio-Visual LocalizationHahyeon Choi, Junhoo Lee, Nojun KwakICCV 2025 · 1 citation
