Self-Supervised Visual Acoustic Matching
Arjun Somayazulu, Changan Chen, Kristen Grauman
摘要
Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both the source and target environments, but this limits the diversity of training data or requires the use of simulated data or heuristics to create paired samples. We propose a self-supervised approach to visual acoustic matching where training samples include only the target scene image and audio-without acoustically mismatched source audio for reference. Our approach jointly learns to disentangle room acoustics and resynthesize audio into the target environment, via a conditional GAN framework and a novel metric that quantifies the level of residual acoustic information in the de-biased audio. Training with either in-the-wild web data or simulated data, we demonstrate it outperforms the state-of-the-art on multiple challenging datasets and a wide variety of real-world audio and environments. Project page: https: //vision.cs.utexas.edu/projects/ss_vam
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu 等NeurIPS 2024 · 被引用 31 次
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 被引用 22 次
- AV-RIR: Audio-Visual Room Impulse Response EstimationAnton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya 等CVPR 2024 · 被引用 15 次
- Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-SpeechRui Liu, Shuwei He, Yifan Hu, Haizhou LiAAAI 2025 · 被引用 8 次
- How Would it Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor ScenesMahnoor Fatima Saad, Ziad Al-HalahICCV 2025 · 被引用 4 次
它引用的顶会 Paper9
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson 等ICML 2020 · 被引用 210 次
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum 等NeurIPS 2022 · 被引用 153 次
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 被引用 80 次
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn 等ICLR 2021 · 被引用 73 次
相关 Paper
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 被引用 42 次
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer 等CVPR 2023
- Audio-Visual Localization by Synthetic Acoustic Image GenerationValentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio MurinoAAAI 2021 · 被引用 9 次
- Sound to Visual Scene Generation by Audio-to-Visual Latent AlignmentSung-Bin Kim, Arda Senocak, Hyunwoo Ha, Andrew Owens 等CVPR 2023
- Cinematic Audio Source Separation Using Visual CuesKang Zhang, Suyeon Lee, Arda Senocak, Joon Son ChungCVPR 2026 · 被引用 1 次
