Global Patch-wise Attention is Masterful Facilitator for Masked Image Modeling
Gongli Xi, Ye Tian, Mengyu Yang, Lanshan Zhang, Xirong Que, Wendong Wang
Abstract
Masked image modeling (MIM), as a self-supervised learning paradigm in computer vision, has gained widespread attention among researchers. MIM operates by training the model to predict masked patches of the image. Given the sparse nature of image semantics, it is imperative to devise a masking strategy that steers the model towards reconstructing high-semantic regions. However, conventional mask strategies often miss these high-semantic regions or lack alignment with the masks and semantics. To solve this, we propose the Global Patch-wise Attention (GPA) framework, a transferable and efficient framework for MIM pre-training. We observe that the attention between patches can be the metric of identifying high-semantic regions, which can guide the model to learn more effective representations. Therefore, we firstly define the global patch-wise attention via vision transformer blocks. Then we design the soft-to-hard mask generation to guide the model gradually focusing on high semantic regions identified by GPA (GPA as a teacher). Finally, we design an extra task to predict GPA (GPA as a feature). Experiments conducted under various settings demonstrate that our proposed GPA framework enables MIM to learn better representations, which benefit the model across a wide range of downstream tasks. Furthermore, our GPA framework can be easily and effectively transferred to various MIM architectures.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b843e5ca-64f6-40e5-871e-a5c3b23f336bCited by top-tier papers1
Ask how each one uses itRelated papers
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang et al.ICML 2023 · 60 citations
- Good Helper Is around You: Attention-Driven Masked Image ModelingZhengqi Liu, Jie Gui, Hao LuoAAAI 2023 · 36 citations
- Stare at What You See: Masked Image Modeling without ReconstructionHongwei Xue, Peng Gao, Hongyang Li, Yu Qiao et al.CVPR 2023
- Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual PretrainingShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023
- Hard Patches Mining for Masked Image ModelingHaochen Wang, Kaiyou Song, Junsong Fan, Yuxi Wang et al.CVPR 2023
