Evolving Losses for Unsupervised Video Representation Learning
A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo
Abstract
We present a new method to learn video representations from large-scale unlabeled video data. Ideally, this representation will be generic and transferable, directly usable for new tasks such as action recognition and zero or few-shot learning. We formulate unsupervised representation learning as a multi-modal, multi-task learning problem, where the representations are shared across different modalities via distillation. Further, we introduce the concept of loss function evolution by using an evolutionary search algorithm to automatically find optimal combination of loss functions capturing many (self-supervised) tasks and modalities. Thirdly, we propose an unsupervised representation evaluation metric using distribution matching to a large unlabeled dataset as a prior constraint, based on Zipf's law. This unsupervised constraint, which is not guided by any labeling, produces similar results to weaklysupervised, task-specific ones. The proposed unsupervised representation learning results in a single RGB network and outperforms previous methods. Notably, it is also more effective than several label-based methods (e.g., ImageNet), with the exception of large, fully labeled video datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fb485e7-e179-4956-aafa-75e65ad502cdCited by top-tier papers45
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
Builds on2
Related papers
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 59 citations
- Improving Video Model Transfer with Dynamic Representation LearningYi Li, Nuno VasconcelosCVPR 2022 · 2 citations
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long et al.AAAI 2021 · 140 citations
- MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeWei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger et al.ICCV 2023 · 52 citations
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
