Self-supervised object detection from audio-visual correspondence
Triantafyllos Afouras, Yuki M. Asano, Francois Fagan, Andrea Vedaldi, Florian Metze
Abstract
We tackle the problem of learning object detectors without supervision. Differently from weakly-supervised object detection, we do not assume image-level class labels. Instead, we extract a supervisory signal from audio-visual data, using the audio component to “teach” the object detector. While this problem is related to sound source localisation, it is considerably harder because the detector must classify the objects by type, enumerate each instance of the object, and do so even when the object is silent. We tackle this problem by first designing a self-supervised framework with a contrastive objective that jointly learns to classify and localise objects. Then, without using any supervision, we simply use these self-supervised labels and boxes to train an image-based object detector. With this, we outperform previous unsupervised and weakly-supervised detectors for the task of object detection and sound source localization. We also show that we can align this detector to ground-truth classes with as little as one label per pseudo-class, and show how our method can learn to detect generic objects that go beyond instruments, such as airplanes and cats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- POP-3D: Open-Vocabulary 3D Occupancy Prediction from ImagesAntonín Vobecký, Oriane Siméoni, David Hurych, Spyridon Gidaris et al.NeurIPS 2023 · 67 citations
- Object-aware Contrastive Learning for Debiased Scene RepresentationSangwoo Mo, Hyunwoo Kang, Kihyuk Sohn, Chun-Liang Li et al.NeurIPS 2021 · 57 citations
- Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and LanguageOtniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep AkataCVPR 2022 · 54 citations
- Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker TrackingYidi Li, Hong Liu, Hao TangAAAI 2022 · 25 citations
- Sound Localization from Motion: Jointly Learning Sound Direction and Camera RotationZiyang Chen, Shengyi Qian, Andrew OwensICCV 2023 · 21 citations
Builds on29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar et al.ICCV 2019 · 901 citations
Related papers
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er et al.AAAI 2023 · 13 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- Weakly-Supervised Audio-Visual SegmentationShentong Mo, Bhiksha RajNeurIPS 2023 · 26 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
