DistInit: Learning Video Representations Without a Single Labeled Video
Rohit Girdhar, Du Tran, Lorenzo Torresani, Deva Ramanan
Abstract
Video recognition models have progressed significantly over the past few years, evolving from shallow classifiers trained on hand-crafted features to deep spatiotemporal networks. However, labeled video data required to train such models has not been able to keep up with the ever increasing depth and sophistication of these networks. In this work we propose an alternative approach to learning video representations that requires no semantically labeled videos, and instead leverages the years of effort in collecting and labeling large and clean still-image datasets. We do so by using state-of-the-art models pre-trained on image datasets as “teachers” to train video models in a distillation framework. We demonstrate that our method learns truly spatiotemporal features, despite being trained only using supervision from still-image networks. Moreover, it learns good representations across different input modalities, using completely uncurated raw video data sources and with different 2D teacher models. Our method obtains strong transfer performance, outperforming standard techniques for bootstrapping video architectures with image based models by 16%. We believe that our approach opens up new approaches for learning spatiotemporal representations from unlabeled video data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
- Representation Learning via Adversarially-Contrastive Optimal TransportAnoop Cherian, Shuchin AeronICML 2020 · 9 citations
- Background Splitting: Finding Rare Classes in a Sea of BackgroundRavi Teja Mullapudi, Fait Poms, William R. Mark, Deva Ramanan et al.CVPR 2021
- NIL: No-data Imitation LearningMert Albaba, Chenhao Li, Markos Diomataris, Omid Taheri et al.CVPR 2026
- Listen to Look: Action Recognition by Previewing AudioRuohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo TorresaniCVPR 2020
Related papers
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2023
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah et al.CVPR 2024 · 3 citations
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
- DreamTeacher: Pretraining Image Backbones with Deep Generative ModelsDaiqing Li, Huan Ling, Amlan Kar, David Acuna et al.ICCV 2023 · 37 citations
- Self-supervised video pretraining yields robust and more human-aligned visual representationsNikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. HénaffNeurIPS 2023 · 27 citations
