Unsupervised Learning From Video With Deep Neural Embeddings
Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, Daniel Yamins
Abstract
Because of the rich dynamical structure of videos and their ubiquity in everyday life, it is a natural idea that video data could serve as a powerful unsupervised learning signal for visual representations. However, instantiating this idea, especially at large scale, has remained a significant artificial intelligence challenge. Here we present the Video Instance Embedding (VIE) framework, which trains deep nonlinear embeddings on video sequence inputs. By learning embedding dimensions that identify and group similar videos together, while pushing inherently different videos apart in the embedding space, VIE captures the strong statistical structure inherent in videos, without the need for external annotation labels. We find that, when trained on a large-scale video dataset, VIE yields powerful representations both for action recognition and single-frame object categorization, showing substantially improving on the state of the art wherever direct comparisons are possible. We show that a twopathway model with both static and dynamic processing pathways is optimal, provide analyses indicating how the model works, and perform ablation studies showing the importance of key architecture and loss function choices. Our results suggest that deep neural embeddings are a promising approach to unsupervised video learning for a wide variety of task domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57524fe1-265a-49df-91e3-0f209322ac76Cited by top-tier papers12
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Spatio-temporal Self-Supervised Representation Learning for 3D Point CloudsSiyuan Huang, Yichen Xie, Song-Chun Zhu, Yixin ZhuICCV 2021 · 259 citations
- Self-supervised learning through the eyes of a childA. Emin Orhan, Vaibhav V. Gupta, Brenden M. LakeNeurIPS 2020 · 119 citations
- SeCo: Exploring Sequence Supervision for Unsupervised Representation LearningTing Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan et al.AAAI 2021 · 118 citations
- Divide and Contrast: Self-supervised Learning from Uncurated DataYonglong Tian, Olivier J. Hénaff, Aäron van den OordICCV 2021 · 112 citations
Builds on3
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 462 citations
Related papers
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
- Vi2CLR: Video and Image for Visual Contrastive Learning of RepresentationAli Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi et al.ICCV 2021 · 65 citations
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 59 citations
- Learning Hierarchical Embedding for Video Instance SegmentationZheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen et al.ACM MM 2021 · 16 citations
