There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge
Francisco Rivera Valverde, Juana Valeria Hurtado, Abhinav Valada
Abstract
Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to localize objects over the image field by solely monitoring the sound in the environment. Thus far, this has only been feasible in scenarios where the camera is static and for single object detection. Moreover, the robustness of these methods has been limited as they primarily rely on RGB images which are highly susceptible to illumination and weather changes. In this work, we present the novel self-supervised MM-DistillNet framework consisting of multiple teachers that leverage diverse modalities including RGB, depth and thermal images, to simultaneously exploit complementary cues and distill knowledge into a single audio student network. We propose the new MTA loss function that facilitates the distillation of information from multimodal teachers in a self-supervised manner. Additionally, we propose a novel self-supervised pretext task for the audio student that enables us to not rely on labor-intensive manual annotations. We introduce a large-scale multimodal dataset with over 113,000 time-synchronized frames of RGB, depth, thermal, and audio modalities. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while being able to detect multiple objects using only sound during inference and even while moving.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f29d005e-2deb-4b9c-9c7b-fbda2361a07cCited by top-tier papers10
- COCOA: Cross Modality Contrastive Learning for Sensor DataShohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V. Smith et al.UbiComp 2022 · 88 citations
- Mix and Localize: Localizing Sound Sources in MixturesXixi Hu, Ziyang Chen, Andrew OwensCVPR 2022 · 50 citations
- Amodal Panoptic SegmentationRohit Mohan, Abhinav ValadaCVPR 2022 · 49 citations
- Boosting 3D Object Detection by Simulating Multimodality on Point CloudsWu Zheng, Mingxuan Hong, Li Jiang, Chi-Wing FuCVPR 2022 · 32 citations
- The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge DistillationZihui Xue, Zhengqi Gao, Sucheng Ren, Hang ZhaoICLR 2023 · 12 citations
Builds on10
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao et al.AAAI 2020 · 199 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
Related papers
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 2 citations
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 3 citations
- XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation LearningPritam Sarkar, Ali EtemadAAAI 2024 · 45 citations
- Anomaly Detection in Video via Self-Supervised and Multi-Task LearningMariana-Iuliana Georgescu, Antonio Barbalau, Radu Tudor Ionescu, Fahad Shahbaz Khan et al.CVPR 2021
- Multi-Task Driven Feature Models for Thermal Infrared TrackingQiao Liu, Xin Li, Zhenyu He, Nana Fan et al.AAAI 2020 · 73 citations
