Masked Image Modeling with Local Multi-Scale Reconstruction
Haoqing Wang, Yehui Tang, Yunhe Wang, Jianyuan Guo, Zhi-Hong Deng, Kai Han
Abstract
Masked Image Modeling (MIM) achieves outstanding success in self-supervised representation learning. Unfortunately, MIM models typically have huge computational burden and slow learning process, which is an inevitable obstacle for their industrial applications. Although the lower layers play the key role in MIM, existing MIM models conduct reconstruction task only at the top layer of encoder. The lower layers are not explicitly guided and the interaction among their patches is only used for calculating new activations. Considering the reconstruction task requires non-trivial inter-patch interactions to reason target signals, we apply it to multiple local layers including lower and upper layers. Further, since the multiple layers expect to learn the information of different scales, we design local multiscale reconstruction, where the lower and upper layers reconstruct fine-scale and coarse-scale supervision signals respectively. This design not only accelerates the representation learning process by explicitly guiding multiple layers, but also facilitates multi-scale semantical understanding to the input. Extensive experiments show that with significantly less pre-training burden, our model achieves comparable or better performance on classification, detection and segmentation tasks than existing MIM models. Code is available with both MindSpore and PyTorch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dbe29c2-8f74-4ef8-aa6f-fd73a227dd81Cited by top-tier papers13
- OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang et al.NeurIPS 2024 · 45 citations
- DropPos: Pre-Training Vision Transformers by Reconstructing Dropped PositionsHaochen Wang, Junsong Fan, Yuxi Wang, Kaiyou Song et al.NeurIPS 2023 · 32 citations
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- Focus Your Attention when Few-Shot ClassificationHaoqing Wang, Shibo Jie, Zhihong DengNeurIPS 2023 · 16 citations
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 12 citations
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- DDAE: Towards Deep Dynamic Vision BERT PretrainingHonghao Chen, Xiangwen Kong, Xiangyu Zhang, Xin Zhao et al.AAAI 2024 · 1 citation
- Good Helper Is around You: Attention-Driven Masked Image ModelingZhengqi Liu, Jie Gui, Hao LuoAAAI 2023 · 36 citations
- Masked Image Modeling with Denoising ContrastKun Yi, Yixiao Ge, Xiaotong Li, Shusheng Yang et al.ICLR 2023 · 8 citations
- Progressively Compressed Auto-Encoder for Self-supervised Representation LearningJin Li, Yaoming Wang, Xiaopeng Zhang, Yabo Chen et al.ICLR 2023
- Region Similarity Representation LearningTete Xiao, Colorado J. Reed, Xiaolong Wang, Kurt Keutzer et al.ICCV 2021 · 128 citations
