Soundify: Matching Sound Effects to Video
David Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi, Nikolas Martelaro
摘要
In the art of video editing, sound helps add character to an object and immerse the viewer within a space. Through formative interviews with professional editors (N=10), we found that the task of adding sounds to video can be challenging. This paper presents Soundify, a system that assists editors in matching sounds to video. Given a video, Soundify identifies matching sounds, synchronizes the sounds to the video, and dynamically adjusts panning and volume to create spatial audio. In a human evaluation study (N=889), we show that Soundify is capable of matching sounds to video out-of-the-box for a diverse range of audio categories. In a within-subjects expert study (N=12), we demonstrate the usefulness of Soundify in helping video editors match sounds to video with lighter workload, reduced task completion time, and improved usability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SonifyAR: Context-Aware Sound Generation in Augmented RealityXia Su, Jon E. Froehlich, Eunyee Koh, Chang XiaoUIST 2024 · 被引用 12 次
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon 等CHI 2026 · 被引用 2 次
- MoSound: An Interactive Tool for Generative Sound Design in Motion GraphicsJialin Huang, Prem Seetharaman, Timothy Richard Langlois, Li-Yi Wei 等CHI 2026 · 被引用 2 次
- Auditorily Embodied Conversational Agents: Effects of Spatialization and Situated Audio Cues on Presence and Social PerceptionYi Fei Cheng, Jarod Bloch, Alexander Wang, Andrea Bianchi 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Audeo: Audio Generation for a Silent Performance VideoKun Su, Xiulong Liu, Eli ShlizermanNeurIPS 2020 · 被引用 78 次
- Crosscast: Adding Visuals to Audio Travel PodcastsHaijun Xia, Jennifer Jacobs, Maneesh AgrawalaUIST 2020 · 被引用 44 次
相关 Paper
- Eventfulness for Interactive Video AlignmentJiatian Sun, Longxiulin Deng, Triantafyllos Afouras, Andrew Owens 等SIGGRAPH 2023 · 被引用 5 次
- D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX MatchingJingyu Liu, Minquan Wang, Ye Ma, Bo Wang 等AAAI 2025 · 被引用 4 次
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song 等ACM MM 2024 · 被引用 6 次
- AutoSFX: Automatic Sound Effect Generation for VideosYujia Wang, Zhongxu Wang, Hua HuangACM MM 2024 · 被引用 2 次
- What's Making That Sound Right Now? Video-Centric Audio-Visual LocalizationHahyeon Choi, Junhoo Lee, Nojun KwakICCV 2025 · 被引用 1 次
