Self-Supervised Moving Vehicle Tracking With Stereo Sound
Chuang Gan, Hang Zhao, Peihao Chen, David D. Cox, Antonio Torralba
Abstract
Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audiovisual data to learn to localize objects (moving vehicles) in a visual reference frame, purely using stereo sound at inference time. Since it is labor-intensive to manually annotate the correspondences between audio and object bounding boxes, we achieve this goal by using the co-occurrence of visual and audio streams in unlabeled videos as a form of self-supervision, without resorting to the collection of ground truth annotations. In particular, we propose a framework that consists of a vision teacher'' network and a stereo-sound student'' network. During training, knowledge embodied in a well-established visual vehicle detection model is transferred to the audio domain using unlabeled videos as a bridge. At test time, the stereo-sound student network can work independently to perform object localization using just stereo audio and camera meta-data, without any visual input. Experimental results on a newly collected Auditory Vehicles Tracking dataset verify that our proposed approach outperforms several baseline approaches. We also demonstrate that our cross-modal auditory localization approach can assist in the visual localization of moving vehicles under poor lighting conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc9b1453-dea9-4175-ba62-5bf8fdf2c9c5Cited by top-tier papers45
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long et al.AAAI 2021 · 140 citations
- Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic SegmentationZhuangwei Zhuang, Rong Li, Kui Jia, Qicheng Wang et al.ICCV 2021 · 129 citations
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360° VideosHeeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee et al.ICCV 2021 · 124 citations
- Beyond the Parts: Learning Multi-view Cross-part Correlation for Vehicle Re-identificationXinchen Liu, Wu Liu, Jinkai Zheng, Chenggang Yan et al.ACM MM 2020 · 97 citations
Builds on1
Related papers
- There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal KnowledgeFrancisco Rivera Valverde, Juana Valeria Hurtado, Abhinav ValadaCVPR 2021
- Self-supervised object detection from audio-visual correspondenceTriantafyllos Afouras, Yuki M. Asano, Francois Fagan, Andrea Vedaldi et al.CVPR 2022 · 50 citations
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er et al.AAAI 2023 · 13 citations
- Supervising Sound Localization by In-the-wild EgomotionAnna Min, Ziyang Chen, Hang Zhao, Andrew OwensCVPR 2025
