Images that Sound: Composing Images and Sounds on a Single Canvas
Ziyang Chen, Daniel Geng, Andrew Owens
摘要
Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these visual spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana 等NeurIPS 2024 · 被引用 39 次
- Stroke of Surprise: Progressive Semantic Illusions in Vector SketchingHuai-Hsun Cheng, Siang-Ling Zhang, Yu-Lun LiuSIGGRAPH 2026 · 被引用 1 次
- LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPascal Chang, Sergio Sancho, Jingwei Tang, Markus Gross 等CVPR 2025
- SpeechOp: Inference-Time Task Composition for Generative Speech ProcessingJustin Lovelace, Rithesh Kumar, Jiaqi Su, Ke Chen 等ICLR 2026
- Video-Guided Foley Sound Generation with Multimodal ControlsZiyang Chen, Prem Seetharaman, Bryan C. Russell, Oriol Nieto 等CVPR 2025
它引用的顶会 Paper68
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Animate and Sound an ImageXihua Wang, Ruihua Song, Chongxuan Li, Xin Cheng 等CVPR 2025
- Sound to Visual Scene Generation by Audio-to-Visual Latent AlignmentSung-Bin Kim, Arda Senocak, Hyunwoo Ha, Andrew Owens 等CVPR 2023
- Generating Realistic Images from In-the-wild SoundsTaegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan KimICCV 2023 · 被引用 11 次
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang 等NeurIPS 2025 · 被引用 8 次
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State SpacesKevin Rojas, Yuchen Zhu, Sichen Zhu, Felix X.-F. Ye 等ICML 2025
