Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation
Ziyang Chen, Shengyi Qian, Andrew Owens
Abstract
The images and sounds that we perceive undergo subtle but geometrically consistent changes as we rotate our heads. In this paper, we use these cues to solve a problem we call Sound Localization from Motion (SLfM): jointly estimating camera rotation and localizing sound sources. We learn to solve these tasks solely through self-supervision. A visual model predicts camera rotation from a pair of images, while an audio model predicts the direction of sound sources from binaural sounds. We train these models to generate predictions that agree with one another. At test time, the models can be deployed independently. To obtain a feature representation that is well-suited to solving this challenging problem, we also propose a method for learning an audio-visual representation through cross-view binauralization: estimating binaural sound from one view, given images and sound from another. Our model can successfully estimate accurate rotations on both real and synthetic scenes, and localize sound sources with accuracy competitive with state-of-the-art self-supervised approaches. Project site: https://ificl.github.io/SLfM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d507484c-7a6d-4191-aa6a-5d85c534fee9Cited by top-tier papers12
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park et al.CVPR 2024 · 47 citations
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 22 citations
- Bioacoustic Geolocation: Species Sounds as Geographic SignalsMustafa Chasmai, Wuao Liu, Subhransu Maji, Grant HornICML 2026 · 5 citations
- Resounding Acoustic Fields with ReciprocityZitong Lan, Yiduo Hao, Mingmin ZhaoNeurIPS 2025 · 3 citations
- Few-shot Acoustic Synthesis with Multimodal Flow MatchingAmandine BrunettoCVPR 2026 · 2 citations
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 551 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
Related papers
- Supervising Sound Localization by In-the-wild EgomotionAnna Min, Ziyang Chen, Hang Zhao, Andrew OwensCVPR 2025
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 27 citations
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh et al.ICCV 2023 · 39 citations
