SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video
Suhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon, Oriol Nieto, Anh Truong
摘要
Sound effects (SFX) are critical to video storytelling by immersing viewers, directing attention, and shaping emotion. However, crafting an effective soundscape is difficult: creators must decide how to source, place, layer, and mix sounds to support the narrative. Generative text-to-SFX tools enable users to create custom sounds, but creators often struggle to describe sounds with words and lack control over individual stems in premixed outputs. We propose SoundStager, an AI-assisted tool for designing generative soundscapes for video. SoundStager analyzes the video narrative to create layered audio scenes (of keynote, signal, soundmark, and archetypal sounds) and supports iterative refinement through a combination of conversational and analog controls. SoundStager’s design was informed by formative studies with six professional sound designers, six video creators, and insights from sound design literature. Our user evaluation with twelve video creators shows that SoundStager enables users to quickly create satisfactory soundscapes while retaining creative control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
- Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative ModelsRyan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry 等CHI 2020 · 被引用 265 次
- AudioGen: Textually Guided Audio GenerationFelix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer 等ICLR 2023 · 被引用 54 次
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 被引用 48 次
相关 Paper
- AutoSFX: Automatic Sound Effect Generation for VideosYujia Wang, Zhongxu Wang, Hua HuangACM MM 2024 · 被引用 2 次
- Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound DesignersPurnima Kamath, Fabio Morreale, Priambudi Lintang Bagaskara, Yize Wei 等CHI 2024 · 被引用 36 次
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin 等CHI 2025 · 被引用 26 次
- Soundify: Matching Sound Effects to VideoDavid Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi 等UIST 2023 · 被引用 15 次
- VidTune: Creating Video Soundtracks with Generative Music and Video-Based ThumbnailsMina Huh, C. Ailie Fraser, Dingzeyu Li, Mira Dontcheva 等CHI 2026 · 被引用 1 次
