SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video
Suhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon, Oriol Nieto, Anh Truong
Abstract
Sound effects (SFX) are critical to video storytelling by immersing viewers, directing attention, and shaping emotion. However, crafting an effective soundscape is difficult: creators must decide how to source, place, layer, and mix sounds to support the narrative. Generative text-to-SFX tools enable users to create custom sounds, but creators often struggle to describe sounds with words and lack control over individual stems in premixed outputs. We propose SoundStager, an AI-assisted tool for designing generative soundscapes for video. SoundStager analyzes the video narrative to create layered audio scenes (of keynote, signal, soundmark, and archetypal sounds) and supports iterative refinement through a combination of conversational and analog controls. SoundStager’s design was informed by formative studies with six professional sound designers, six video creators, and insights from sound design literature. Our user evaluation with twelve video creators shows that SoundStager enables users to quickly create satisfactory soundscapes while retaining creative control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a36440ea-5c94-4d73-9161-1e8e832f3dadBuilds on21
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
- Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative ModelsRyan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry et al.CHI 2020 · 265 citations
- AudioGen: Textually Guided Audio GenerationFelix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer et al.ICLR 2023 · 54 citations
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 48 citations
Related papers
- AutoSFX: Automatic Sound Effect Generation for VideosYujia Wang, Zhongxu Wang, Hua HuangACM MM 2024 · 2 citations
- Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound DesignersPurnima Kamath, Fabio Morreale, Priambudi Lintang Bagaskara, Yize Wei et al.CHI 2024 · 36 citations
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin et al.CHI 2025 · 26 citations
- Soundify: Matching Sound Effects to VideoDavid Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi et al.UIST 2023 · 15 citations
- VidTune: Creating Video Soundtracks with Generative Music and Video-Based ThumbnailsMina Huh, C. Ailie Fraser, Dingzeyu Li, Mira Dontcheva et al.CHI 2026 · 1 citation
