Telling Left From Right: Learning Spatial Correspondence of Sight and Sound
Karren Yang, Bryan C. Russell, Justin Salamon
Abstract
Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the sensory streams. We propose a novel self-supervised task to leverage an orthogonal principle: matching spatial information in the audio stream to the positions of sound sources in the visual stream. Our approach is simple yet effective. We train a model to determine whether the left and right audio channels have been flipped, forcing it to reason about spatial localization across the visual and audio streams. To train and evaluate our method, we introduce a large-scale video dataset, YouTube-ASMR-300K, with spatial audio comprising over 900 hours of footage. We demonstrate that understanding spatial correspondence enables models to perform better on three audiovisual tasks, achieving quantitative gains over supervised and self-supervised baselines that do not leverage spatial audio cues. We also show how to extend our self-supervised approach to 360 degree videos with ambisonic audio. * Work done at Adobe Research during KYs summer internship. 1 These videos are provided in the Supplementary Materials. We encourage you to watch and listen to the videos wearing headphones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d94e2d8a-94ac-43b7-becc-aa2fb8915792Cited by top-tier papers20
- Active Contrastive Learning of Audio-Visual Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongICLR 2021 · 109 citations
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
- ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation LearningSangho Lee, Jiwan Chung, Youngjae Yu, Gunhee Kim et al.ICCV 2021 · 73 citations
- Mix and Localize: Localizing Sound Sources in MixturesXixi Hu, Ziyang Chen, Andrew OwensCVPR 2022 · 50 citations
- Contrastive Learning of Global and Local Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongNeurIPS 2021 · 45 citations
Builds on3
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
Related papers
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 149 citations
- Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio GenerationYan-Bo Lin, Yu-Chiang Frank WangAAAI 2021 · 25 citations
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 3 citations
- CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentEdson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati et al.CVPR 2025
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
