SoundBrush: Sound as a Brush for Visual Scene Editing
Sung-Bin Kim, Kim Jun-Seong, Junseok Ko, Yewon Kim, Tae-Hyun Oh
Abstract
We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorporate audio information for editing visual scenes. Inspired by existing image-editing works, we frame this task as a supervised learning problem and leverage various off-the-shelf models to construct a sound-paired visual scene editing dataset for training. This richly generated dataset enables SoundBrush to learn to map audio features into the textual space of the LDM, allowing for visual scene editing guided by diverse in-the-wild sound. Unlike existing methods, SoundBrush can accurately manipulate the overall scenery or even insert sounding objects to best match the input sound semantics while preserving the original content. Furthermore, by integrating with novel view synthesis techniques, our framework can be extended to edit 3D scenes, facilitating sound-driven 3D scene manipulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ec4c873-4385-4b5f-a664-c49446167316Cited by top-tier papers3
- SounDiT: Geo-Contextual Soundscape-to-Landscape GenerationJunbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang et al.CVPR 2026 · 3 citations
- CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image GenerationHyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim et al.ACM MM 2025
- LandCraft: Designing the Structured 3D Landscapes via Text GuidanceZhihao Liu, Fang Liu, Weihao Xuan, Naoto YokoyaAAAI 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- AUDIT: Audio Editing by Following Instructions with Latent Diffusion ModelsYuancheng Wang, Zeqian Ju, Xu Tan, Lei He et al.NeurIPS 2023 · 120 citations
- Sound-Guided Semantic Image ManipulationSeung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon et al.CVPR 2022 · 43 citations
- Customized Condition Controllable Generation for Video SoundtrackFan Qi, Kunsheng Ma, Changsheng XuCVPR 2025
- ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image ManipulationYasheng Sun, Yifan Yang, Houwen Peng, Yifei Shen et al.NeurIPS 2023 · 71 citations
- MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video GenerationMingzhen Sun, Weining Wang, Yanyuan Qiao, Jiahui Sun et al.ACM MM 2024 · 4 citations
