PooDLe🐩: Pooled and dense self-supervised learning from naturalistic videos
Alex N. Wang, Christopher Hoang, Yuwen Xiong, Yann LeCun, Mengye Ren
摘要
Self-supervised learning has driven significant progress in learning from singlesubject, iconic images. However, there are still unanswered questions about the use of minimally-curated, naturalistic video data, which contain dense scenes with many independent objects, imbalanced class distributions, and varying object sizes. In this paper, we propose PooDLe, a self-supervised learning method that combines an invariance-based objective on pooled representations with a dense SSL objective that enforces equivariance to optical flow warping. Our results show that a unified objective applied at multiple feature scales is essential for learning effective image representations from naturalistic videos. We validate our method with experiments on the BDD100K driving video dataset and the Walking Tours first-person video dataset, demonstrating its ability to capture spatial understanding from a dense objective and semantic understanding via a pooled representation objective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 被引用 4 次
- Midway Network: Learning Representations for Recognition and Motion from Latent DynamicsChristopher Hoang, Mengye RenICLR 2026 · 被引用 2 次
- Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric VideoYuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li 等CVPR 2026 · 被引用 1 次
- Featurising Pixels from Dynamic 3D Scenes with Linear In-Context LearnersNikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki, David Joseph Tan 等CVPR 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Self-Supervised Representation Learning from Flow EquivarianceYuwen Xiong, Mengye Ren, Wenyuan Zeng, Raquel Urtasun WaabiICCV 2021 · 被引用 32 次
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset BiasesSenthil Purushwalkam, Abhinav GuptaNeurIPS 2020 · 被引用 240 次
- Shepherding Slots to Objects: Towards Stable and Robust Object-Centric LearningJinwoo Kim, Janghyuk Choi, Ho-Jin Choi, Seon Joo KimCVPR 2023
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li 等CVPR 2022 · 被引用 29 次
- SCALOR: Generative World Models with Scalable Object RepresentationsJindong Jiang, Sepehr Janghorbani, Gerard de Melo, Sungjin AhnICLR 2020 · 被引用 152 次
