Stochastic positional embeddings improve masked image modeling
Amir Bar, Florian Bordes, Assaf Shocher, Mido Assran, Pascal Vincent, Nicolas Ballas, Trevor Darrell, Amir Globerson, Yann LeCun
摘要
Masked Image Modeling (MIM) is a promising self-supervised learning approach that enables learning from unlabeled images. Despite its recent success, learning good representations through MIM remains challenging because it requires predicting the right semantic content in accurate locations. For example, given an incomplete picture of a dog, we can guess that there is a tail, but we cannot determine its exact location. In this work, we propose to incorporate location uncertainty into MIM by using stochastic positional embeddings (StoP). Specifically, we condition the model on stochastic masked token positions drawn from a Gaussian distribution. StoP reduces overfitting to location features and guides the model toward learning features that are more robust to location uncertainties. Quantitatively, StoP improves downstream MIM performance on a variety of downstream tasks, including on ImageNet linear probing using ViT-B, and for ViT-H using of the data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Scaling Language-Free Visual Representation LearningDavid Fan, Shengbang Tong, Jiachen Zhu, Koustuv Sinha 等ICCV 2025 · 被引用 4 次
- From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer LearningYang Liu, Qianqian Xu, Peisong Wen, Siran Dai 等CVPR 2026 · 被引用 2 次
- Text-Conditional JEPA for Learning Semantically Rich Visual RepresentationsChen Huang, Xianhang Li, Vimal Thilak, Etai Littwin 等ICML 2026 · 被引用 1 次
- Sample- and Parameter-Efficient Auto-Regressive Image ModelsElad Amrani, Leonid Karlinsky, Alex M. BronsteinCVPR 2025
- When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation LearningYang Liu, Qianqian Xu, Peisong Wen, Siran Dai 等CVPR 2025
它引用的顶会 Paper23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Representation Learning by Detecting Incorrect Location EmbeddingsSepehr Sameni, Simon Jenni, Paolo FavaroAAAI 2023 · 被引用 8 次
- DDAE: Towards Deep Dynamic Vision BERT PretrainingHonghao Chen, Xiangwen Kong, Xiangyu Zhang, Xin Zhao 等AAAI 2024 · 被引用 1 次
- Beyond [cls]: Exploring the True Potential of Masked Image Modeling RepresentationsMarcin Przewiezlikowski, Randall Balestriero, Wojciech Jasinski, Marek Smieja 等ICCV 2025 · 被引用 4 次
- Positional Label for Self-Supervised Vision TransformerZhemin Zhang, Xun GongAAAI 2023 · 被引用 12 次
- Pre-training with Random Orthogonal Projection Image ModelingMaryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr KoniuszICLR 2024 · 被引用 15 次
