Sound-Guided Semantic Image Manipulation
Seung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon, Chanyoung Kim, Jinkyu Kim, Sangpil Kim
Abstract
The recent success of the generative model shows that leveraging the multi-modal embedding space can manipu-late an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy due to the dynamic characteristics of the sources. Especially, sound can convey vivid emotions and dynamic expressions of the real world. Here, we propose a framework that directly encodes sound into the multi-modal (image-text) embedding space and manipulates an image from the space. Our audio encoder is trained to pro-duce a latent representation from an audio input, which is forced to be aligned with image and text representations in the multi-modal embedding space. We use a direct latent op-timization method based on aligned embeddings for sound-guided image manipulation. We also show that our method can mix different modalities, i.e., text and audio, which en-rich the variety of the image modification. The experiments on zero-shot audio classification and semantic-level image classification show that our proposed model outperforms other text and sound-guided state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc30da82-8e84-4e7c-9ebf-b1bd7e144e0fCited by top-tier papers12
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park et al.CVPR 2024 · 47 citations
- The Power of Sound (TPoS): Audio Reactive Video Generation with Stable DiffusionYujin Jeong, Wonjeong Ryoo, Seunghyun Lee, Dabin Seo et al.ICCV 2023 · 41 citations
- GlueGen: Plug and Play Multi-modal Encoders for X-to-image GenerationCan Qin, Ning Yu, Chen Xing, Shu Zhang et al.ICCV 2023 · 27 citations
- Generating Realistic Images from In-the-wild SoundsTaegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan KimICCV 2023 · 11 citations
- SoundBrush: Sound as a Brush for Visual Scene EditingSung-Bin Kim, Kim Jun-Seong, Junseok Ko, Yewon Kim et al.AAAI 2025 · 4 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Designing an encoder for StyleGAN image manipulationOmer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik et al.SIGGRAPH 2021 · 692 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
Related papers
- Sound to Visual Scene Generation by Audio-to-Visual Latent AlignmentSung-Bin Kim, Arda Senocak, Hyunwoo Ha, Andrew Owens et al.CVPR 2023
- High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningChao Xu, Junwei Zhu, Jiangning Zhang, Yue Han et al.CVPR 2023
- BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic DataXuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang et al.ACM MM 2023 · 12 citations
- Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersYazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang et al.CVPR 2024 · 25 citations
- Data-Efficient Multimodal Fusion on a Single GPUNoël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze et al.CVPR 2024 · 6 citations
