Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation
Yan-Bo Lin, Yu-Chiang Frank Wang
Abstract
Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which would degrade the user experience due to the lack of ambient information. To address this issue, we propose an audio spatialization framework to convert a monaural video into a binaural one exploiting the relationship across audio and visual components. By preserving the left-right consistency in both audio and visual modalities, our learning strategy can be viewed as a self-supervised learning technique, and alleviates the dependency on a large amount of video data with ground truth binaural audio data during training. Experiments on benchmark datasets confirm the effectiveness of our proposed framework in both semi-supervised and fully supervised scenarios, with ablation studies and visualization further support the use of our model for audio spatialization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9b40440-0470-4e15-b16c-ce3dc26df1f8Cited by top-tier papers5
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 27 citations
- AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic SynthesisSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng et al.NeurIPS 2024 · 22 citations
- Sound Localization from Motion: Jointly Learning Sound Direction and Camera RotationZiyang Chen, Shengyi Qian, Andrew OwensICCV 2023 · 21 citations
- Cyclic Learning for Binaural Audio Generation and LocalizationZhaojian Li, Bin Zhao, Yuan YuanCVPR 2024
Builds on5
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 95 citations
- Vision-Infused Deep Audio InpaintingHang Zhou, Ziwei Liu, Xudong Xu, Ping Luo et al.ICCV 2019 · 92 citations
- Music Gesture for Visual Sound SeparationChuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum et al.CVPR 2020
Related papers
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 27 citations
- Telling Left From Right: Learning Spatial Correspondence of Sight and SoundKarren Yang, Bryan C. Russell, Justin SalamonCVPR 2020
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 3 citations
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 3 citations
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 4 citations
