Generating Realistic Images from In-the-wild Sounds
Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan Kim
Abstract
Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous studies have focused on generating images from sound in limited categories or music. In this paper, we propose a novel approach to generate images from in-the-wild sounds. First, we convert sound into text using audio captioning. Second, we propose audio attention and sentence attention to represent the rich characteristics of sound and visualize the sound. Lastly, we propose a direct sound optimization with CLIPscore and AudioCLIP and generate images with a diffusion-based model. In experiments, it shows that our model is able to generate high quality images from wild sounds and outperforms baselines in both quantitative and qualitative evaluations on wild audio datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4eb5384-8081-4a63-bb6c-9b4f299e09eeCited by top-tier papers3
- Audio-sync Video Instance Editing with Granularity-Aware Mask RefinerHaojie Zheng, Shuchen Weng, Jingqi Liu, Siqi Yang et al.CVPR 2026 · 8 citations
- MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic AlignmentHao Zhou, Xiaobao Guo, Yuzhe Zhu, Adams Wai-Kin KongAAAI 2026 · 2 citations
- CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image GenerationHyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim et al.ACM MM 2025
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic DataSreyan Ghosh, Sonal Kumar, Zhifeng Kong, Rafael Valle et al.ICLR 2025
- Sound-Guided Semantic Image ManipulationSeung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon et al.CVPR 2022 · 43 citations
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang et al.ICML 2026 · 35 citations
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 42 citations
- Sound to Visual Scene Generation by Audio-to-Visual Latent AlignmentSung-Bin Kim, Arda Senocak, Hyunwoo Ha, Andrew Owens et al.CVPR 2023
