Audio-Visual Localization by Synthetic Acoustic Image Generation
Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino
摘要
Acoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphones. However, acoustic images are typically generated by cumbersome microphone arrays, which are not as widespread as ordinary microphones mounted on optical cameras. To exploit this empowered modality while using standard microphones and cameras we propose to leverage the generation of synthetic acoustic images from common audio-video data for the task of audio-visual localization. The generation of synthetic acoustic images is obtained by a novel deep architecture, based on Variational Autoencoder and U-Net models, which is trained to reconstruct the ground truth spatialized audio data collected by a microphone array, from the associated video and its corresponding monaural audio signal. Namely, the model learns how to mimic what an array of microphones can produce in the same conditions. We assess the quality of the generated synthetic acoustic images on the task of unsupervised sound source localization in a qualitative and quantitative manner, while also considering standard generation metrics. Our model is evaluated by considering both multimodal datasets containing acoustic images, used for the training, and unseen datasets containing just monaural audio signals and RGB frames, showing to reach more accurate localization results as compared to the state of the art. VAE VAE bn AE VAE mse VAE huber VAE 2s VAE 0s MSE 1.1426±0.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox 等ICCV 2019 · 被引用 157 次
- Associative Variational Auto-Encoder with Distributed Latent Spaces and AssociatorsDae Ung Jo, Byeongju Lee, Jongwon Choi, Haanju Yoo 等AAAI 2020 · 被引用 8 次
- Music Gesture for Visual Sound SeparationChuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum 等CVPR 2020
相关 Paper
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong 等ICCV 2023 · 被引用 57 次
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 被引用 32 次
- Sound to Visual Scene Generation by Audio-to-Visual Latent AlignmentSung-Bin Kim, Arda Senocak, Hyunwoo Ha, Andrew Owens 等CVPR 2023
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 被引用 42 次
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 被引用 27 次
