Self-supervised video pretraining yields robust and more human-aligned visual representations
Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. Hénaff
Abstract
Humans learn powerful representations of objects and scenes by observing how they evolve over time. Yet, outside of specific tasks that require explicit temporal understanding, static image pretraining remains the dominant paradigm for learning visual foundation models. We question this mismatch, and ask whether video pretraining can yield visual representations that bear the hallmarks of human perception: generalisation across tasks, robustness to perturbations, and consistency with human judgements. To that end we propose a novel procedure for curating videos, and develop a contrastive framework which learns from the complex transformations therein. This simple paradigm for distilling knowledge from videos, called VITO, yields general representations that far outperform prior video pretraining methods on image understanding tasks, and image pretraining methods on video understanding tasks. Moreover, VITO representations are significantly more robust to natural and synthetic deformations than image-, video-, and adversarially-trained ones. Finally, VITO's predictions are strongly aligned with human judgements, surpassing models that were specifically trained for that purpose. Together, these results suggest that video pretraining could be a simple way of learning unified, robust, and human-aligned representations of the visual world.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c32c437-8e06-45cf-b648-2622607404edCited by top-tier papers11
- Learning predictable and robust neural representations by straightening image sequencesXueyan Niu, Cristina Savin, Eero P. SimoncelliNeurIPS 2024 · 13 citations
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 4 citations
- LayerLock: Non-Collapsing Representation Learning with Progressive FreezingGoker Erdogan, Nikhil Parthasarathy, Catalin Ionescu, Drew A. Hudson et al.ICCV 2025 · 3 citations
- Trackverse: a Large-Scale Object-Centric Video Dataset for Image-Level Representation LearningYibing Wei, Samuel Church, Victor Suciu, Jinhong Lin et al.ICCV 2025 · 2 citations
- Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of ClassifiersThomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann et al.NeurIPS 2025 · 2 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 59 citations
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani et al.ICLR 2023 · 35 citations
- Video OWL-ViT: Temporally-consistent open-world localization in videoGeorg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic et al.ICCV 2023 · 22 citations
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 14 citations
