Cyclic Learning for Binaural Audio Generation and Localization
Zhaojian Li, Bin Zhao, Yuan Yuan
Abstract
Binaural audio is obtained by simulating the biological structure of human ears, which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to synthesize binaural audio, thereby avoiding expensive binaural audio recording. However, most existing methods directly use the entire scene as a guide, ignoring the correspondence between sounds and sounding objects. In this paper, we advocate generating binaural audio using finegrained raw waveform and object-level visual information as guidance. Specifically, we propose a Cyclic Locatingand-UPmixing (CLUP) framework that jointly learns visual sounding object localization and binaural audio generation. Visual sounding object localization establishes the correspondence between specific visual objects and sound modalities, which provides object-aware guidance to improve binaural generation performance. Meanwhile, the spatial information contained in the generated binaural audio can further improve the performance of sounding object localization. In this case, visual sounding object localization and binaural audio generation can achieve cyclic learning and benefit from each other. Experimental results demonstrate that on the FAIR-Play benchmark dataset, our method is significantly ahead of the existing baselines in multiple evaluation metrics (STFT↓: 0.787 vs. 0.851, ENV↓: 0.128 vs. 0.134, WAV↓: 5.244 vs. 5.684, SNR↑: 7.546 vs. 7.044).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 338f6b05-0c36-4192-8cc5-60fa69991e47Cited by top-tier papers7
- Multimodal Neural Acoustic Fields for Immersive Mixed RealityGuaneen Tong, Johnathan Chi-Ho Leung, Xi Peng, Haosheng Shi et al.IEEE VR 2025 · 4 citations
- Gotta Hear Them All: Towards Sound Source Aware Audio GenerationWei Guo, Heng Wang, Jianbo Ma, Weidong CaiAAAI 2026 · 2 citations
- ISDrama: Immersive Spatial Drama Generation through Multimodal PromptingYu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu et al.ACM MM 2025 · 1 citation
- CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio GenerationYuanhong Chen, Kazuki Shimada, Christian Simon, Yukara Ikemiya et al.ACM MM 2025 · 1 citation
- ViSAGe: Video-to-Spatial Audio GenerationJaeyeon Kim, Heeseung Yun, Gunhee KimICLR 2025
Builds on17
- Class Re-Activation Maps for Weakly-Supervised Semantic SegmentationZhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua et al.CVPR 2022 · 223 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 92 citations
- BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio SynthesisYichong Leng, Zehua Chen, Junliang Guo, Haohe Liu et al.NeurIPS 2022 · 86 citations
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICLR 2021 · 73 citations
Related papers
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound SeparationYapeng Tian, Di Hu, Chenliang XuCVPR 2021
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 32 citations
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 27 citations
- In-the-wild Audio Spatialization with Flexible Text-guided LocalizationTianrui Pan, Jie Liu, Zewen Huang, Jie Tang et al.ACL 2025 · 2 citations
- TAS: Personalized Text-guided Audio SpatializationZhaojian Li, Bin Zhao, Yuan YuanACM MM 2024 · 4 citations
