Labelling unlabelled videos from scratch with multi-modal self-supervision
Yuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea Vedaldi
Abstract
A large part of the current success of deep learning lies in the effectiveness of data -more precisely: labelled data. Yet, labelling a dataset with human annotation continues to carry high costs, especially for videos. While in the image domain, recent methods have allowed to generate meaningful (pseudo-) labels for unlabelled datasets without supervision, this development is missing for the video domain where learning feature representations is the current focus. In this work, we a) show that unsupervised labelling of a video dataset does not come for free from strong feature encoders and b) propose a novel clustering method that allows pseudo-labelling of a video dataset without any human annotations, by leveraging the natural correspondence between the audio and visual modalities. An extensive analysis shows that the resulting clusters have high semantic overlap to ground truth human labels. We further introduce the first benchmarking results on unsupervised labelling of common video datasets Kinetics, Kinetics-Sound, VGG-Sound and AVE 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b68460b-aae0-40ab-be77-97f6ff473e94Cited by top-tier papers54
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel et al.NeurIPS 2020 · 805 citations
- Multi-level Feature Learning for Contrastive Multi-view ClusteringJie Xu, Huayi Tang, Yazhou Ren, Liang Peng et al.CVPR 2022 · 335 citations
- Unsupervised Semantic Segmentation by Contrasting Object Mask ProposalsWouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Luc Van GoolICCV 2021 · 285 citations
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze et al.ICLR 2021 · 269 citations
- Leveraging Real Talking Faces via Self-Supervision for Robust Forgery DetectionAlexandros Haliassos, Rodrigo Mira, Stavros Petridis, Maja PanticCVPR 2022 · 138 citations
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Prototypical Contrastive Learning of Unsupervised RepresentationsJunnan Li, Pan Zhou, Caiming Xiong, Steven C. H. HoiICLR 2021 · 484 citations
Related papers
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er et al.AAAI 2023 · 13 citations
- Multiview Pseudo-Labeling for Semi-supervised Learning from VideoBo Xiong, Haoqi Fan, Kristen Grauman, Christoph FeichtenhoferICCV 2021 · 54 citations
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
- XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation LearningPritam Sarkar, Ali EtemadAAAI 2024 · 45 citations
