Self-supervised video pretraining yields robust and more human-aligned visual representations
Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. Hénaff
摘要
Humans learn powerful representations of objects and scenes by observing how they evolve over time. Yet, outside of specific tasks that require explicit temporal understanding, static image pretraining remains the dominant paradigm for learning visual foundation models. We question this mismatch, and ask whether video pretraining can yield visual representations that bear the hallmarks of human perception: generalisation across tasks, robustness to perturbations, and consistency with human judgements. To that end we propose a novel procedure for curating videos, and develop a contrastive framework which learns from the complex transformations therein. This simple paradigm for distilling knowledge from videos, called VITO, yields general representations that far outperform prior video pretraining methods on image understanding tasks, and image pretraining methods on video understanding tasks. Moreover, VITO representations are significantly more robust to natural and synthetic deformations than image-, video-, and adversarially-trained ones. Finally, VITO's predictions are strongly aligned with human judgements, surpassing models that were specifically trained for that purpose. Together, these results suggest that video pretraining could be a simple way of learning unified, robust, and human-aligned representations of the visual world.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Learning predictable and robust neural representations by straightening image sequencesXueyan Niu, Cristina Savin, Eero P. SimoncelliNeurIPS 2024 · 被引用 13 次
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 被引用 4 次
- LayerLock: Non-Collapsing Representation Learning with Progressive FreezingGoker Erdogan, Nikhil Parthasarathy, Catalin Ionescu, Drew A. Hudson 等ICCV 2025 · 被引用 3 次
- Trackverse: a Large-Scale Object-Centric Video Dataset for Image-Level Representation LearningYibing Wei, Samuel Church, Victor Suciu, Jinhong Lin 等ICCV 2025 · 被引用 2 次
- Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of ClassifiersThomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang 等ICCV 2023 · 被引用 266 次
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 被引用 59 次
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani 等ICLR 2023 · 被引用 35 次
- Video OWL-ViT: Temporally-consistent open-world localization in videoGeorg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic 等ICCV 2023 · 被引用 22 次
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 被引用 14 次
