Hard Patches Mining for Masked Image Modeling
Haochen Wang, Kaiyou Song, Junsong Fan, Yuxi Wang, Jin Xie, Zhaoxiang Zhang
摘要
Masked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask strategies. Intuitively, this procedure can be considered as training a student (the model) on solving given problems (predict masked patches). However, we argue that the model should not only focus on solving given problems, but also stand in the shoes of a teacher to produce a more challenging problem by itself. To this end, we propose Hard Patches Mining (HPM), a brand-new framework for MIM pre-training. We observe that the reconstruction loss can naturally be the metric of the difficulty of the pretraining task. Therefore, we introduce an auxiliary loss predictor, predicting patch-wise losses first and deciding where to mask next. It adopts a relative relationship learning strategy to prevent overfitting to exact reconstruction loss values. Experiments under various settings demonstrate the effectiveness of HPM in constructing masked images. Furthermore, we empirically find that solely introducing the loss prediction objective leads to powerful representations, verifying the efficacy of the ability to be aware of where is hard to reconstruct. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionJunkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang 等NeurIPS 2023 · 被引用 46 次
- DropPos: Pre-Training Vision Transformers by Reconstructing Dropped PositionsHaochen Wang, Junsong Fan, Yuxi Wang, Kaiyou Song 等NeurIPS 2023 · 被引用 32 次
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time SeriesSimon A. Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee 等ICLR 2026 · 被引用 22 次
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 被引用 18 次
- Multi-view Masked Contrastive Representation Learning for Endoscopic Video AnalysisKai Hu, Ye Xiao, Yuan Zhang, Xieping GaoNeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- Global Patch-wise Attention is Masterful Facilitator for Masked Image ModelingGongli Xi, Ye Tian, Mengyu Yang, Lanshan Zhang 等ACM MM 2024 · 被引用 1 次
- Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual PretrainingShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023
- Stare at What You See: Masked Image Modeling without ReconstructionHongwei Xue, Peng Gao, Hongyang Li, Yu Qiao 等CVPR 2023
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
- Patch-Aware Sample Selection for Efficient Masked Image ModelingZhengyang Zhuge, Jiaxing Wang, Yong Li, Yongjun Bao 等AAAI 2024 · 被引用 4 次
