Understanding Masked Image Modeling via Learning Occlusion Invariant Feature
Xiangwen Kong, Xiangyu Zhang
摘要
Recently, Masked Image Modeling (MIM) achieves great success in self-supervised visual recognition. However, as a reconstruction-based framework, it is still an open question to understand how MIM works, since MIM appears very different from previous well-studied siamese approaches such as contrastive learning. In this paper, we propose a new viewpoint: MIM implicitly learns occlusion-invariant features, which is analogous to other siamese methods while the latter learns other invariance. By relaxing MIM formulation into an equivalent siamese form, MIM methods can be interpreted in a unified framework with conventional methods, among which only a) data transformations, i.e. what invariance to learn, and b) similarity measurements are different. Furthermore, taking MAE (He et al., 2021) as a representative example of MIM, we empirically find the success of MIM models relates a little to the choice of similarity functions, but the learned occlusion invariant feature introduced by masked image -- it turns out to be a favored initialization for vision transformers, even though the learned feature could be less semantic. We hope our findings could inspire researchers to develop more powerful self-supervised methods in computer vision community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- EEGPT: Pretrained Transformer for Universal and Reliable Representation of EEG SignalsGuangyu Wang, Wenchao Liu, Yuhong He, Cong Xu 等NeurIPS 2024 · 被引用 267 次
- Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative PretrainingZekun Qi, Runpei Dong, Guofan Fan, Zheng Ge 等ICML 2023 · 被引用 209 次
- ReMasker: Imputing Tabular Data with Masked AutoencodingTianyu Du, Luca Melis, Ting WangICLR 2024 · 被引用 41 次
- MAPSeg: Unified Unsupervised Domain Adaptation for Heterogeneous Medical Image Segmentation Based on 3D Masked Autoencoding and Pseudo-LabelingXuzhe Zhang, Yuhao Wu, Elsa D. Angelini, Ang Li 等CVPR 2024 · 被引用 24 次
- RevColV2: Exploring Disentangled Representations in Masked Image ModelingQi Han, Yuxuan Cai, Xiangyu ZhangNeurIPS 2023 · 被引用 16 次
它引用的顶会 Paper26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
- Adversarial Masking for Self-Supervised LearningYuge Shi, N. Siddharth, Philip H. S. Torr, Adam R. KosiorekICML 2022 · 被引用 110 次
- Masked Distillation Advances Self-Supervised Transformer Architecture SearchCaixia Yan, Xiaojun Chang, Zhihui Li, Lina Yao 等ICLR 2024 · 被引用 3 次
- Masked Image Modeling with Denoising ContrastKun Yi, Yixiao Ge, Xiaotong Li, Shusheng Yang 等ICLR 2023 · 被引用 8 次
- Siamese Image Modeling for Self-Supervised Vision Representation LearningChenxin Tao, Xizhou Zhu, Weijie Su, Gao Huang 等CVPR 2023
