Hard Patches Mining for Masked Image Modeling
Haochen Wang, Kaiyou Song, Junsong Fan, Yuxi Wang, Jin Xie, Zhaoxiang Zhang
Abstract
Masked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask strategies. Intuitively, this procedure can be considered as training a student (the model) on solving given problems (predict masked patches). However, we argue that the model should not only focus on solving given problems, but also stand in the shoes of a teacher to produce a more challenging problem by itself. To this end, we propose Hard Patches Mining (HPM), a brand-new framework for MIM pre-training. We observe that the reconstruction loss can naturally be the metric of the difficulty of the pretraining task. Therefore, we introduce an auxiliary loss predictor, predicting patch-wise losses first and deciding where to mask next. It adopts a relative relationship learning strategy to prevent overfitting to exact reconstruction loss values. Experiments under various settings demonstrate the effectiveness of HPM in constructing masked images. Furthermore, we empirically find that solely introducing the loss prediction objective leads to powerful representations, verifying the efficacy of the ability to be aware of where is hard to reconstruct. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d6a80ae-9a42-4642-9001-b9b477da50b0Cited by top-tier papers29
- HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionJunkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang et al.NeurIPS 2023 · 46 citations
- DropPos: Pre-Training Vision Transformers by Reconstructing Dropped PositionsHaochen Wang, Junsong Fan, Yuxi Wang, Kaiyou Song et al.NeurIPS 2023 · 32 citations
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time SeriesSimon A. Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee et al.ICLR 2026 · 22 citations
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- Multi-view Masked Contrastive Representation Learning for Endoscopic Video AnalysisKai Hu, Ye Xiao, Yuan Zhang, Xieping GaoNeurIPS 2024 · 13 citations
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Global Patch-wise Attention is Masterful Facilitator for Masked Image ModelingGongli Xi, Ye Tian, Mengyu Yang, Lanshan Zhang et al.ACM MM 2024 · 1 citation
- Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual PretrainingShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023
- Stare at What You See: Masked Image Modeling without ReconstructionHongwei Xue, Peng Gao, Hongyang Li, Yu Qiao et al.CVPR 2023
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang et al.ICML 2023 · 60 citations
- Patch-Aware Sample Selection for Efficient Masked Image ModelingZhengyang Zhuge, Jiaxing Wang, Yong Li, Yongjun Bao et al.AAAI 2024 · 4 citations
