Towards Understanding Why Mask Reconstruction Pretraining Helps in Downstream Tasks
Jiachun Pan, Pan Zhou, Shuicheng Yan
摘要
For unsupervised pretraining, mask-reconstruction pretraining (MRP) approaches, e.g. MAE (He et al., 2021) and data2vec (Baevski et al., 2022) , randomly mask input patches and then reconstruct the pixels or semantic features of these masked patches via an auto-encoder. Then for a downstream task, supervised fine-tuning the pretrained encoder remarkably surpasses the conventional "supervised learning" (SL) trained from scratch. However, it is still unclear 1) how MRP performs semantic feature learning in the pretraining phase and 2) why it helps in downstream tasks. To solve these problems, we first theoretically show that on an auto-encoder of a two/one-layered convolution encoder/decoder, MRP can capture all discriminative features of each potential semantic class in the pretraining dataset. Then considering the fact that the pretraining dataset is of huge size and high diversity and thus covers most features in downstream dataset, in fine-tuning phase, the pretrained encoder can capture as much features as it can in downstream datasets, and would not lost these features with theoretical guarantees. In contrast, SL only randomly captures some features due to lottery ticket hypothesis. So MRP provably achieves better performance than SL on the classification tasks. Experimental results testify to our data assumptions and also our theoretical implications. * Equal contribution. Pan Jiachun did this work during an internship at Sea AI Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- How Mask Matters: Towards Theoretical Understandings of Masked AutoencodersQi Zhang, Yifei Wang, Yisen WangNeurIPS 2022 · 被引用 119 次
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 被引用 18 次
- Understanding and Enhancing Mask-Based Pretraining towards Universal RepresentationsMingze Dong, Leda Wang, Yuval KlugerNeurIPS 2025 · 被引用 3 次
- Towards Understanding Why FixMatch Generalizes Better Than Supervised LearningJingyang Li, Jiachun Pan, Vincent Y. F. Tan, Kim-Chuan Toh 等ICLR 2025
- Learning Mask Invariant Mutual Information for Masked Image ModelingTao Huang, Yanxiang Ma, Shan You, Chang XuICLR 2025
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- Masked Feature Prediction for Self-Supervised Visual Pre-TrainingChen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu 等CVPR 2022 · 被引用 524 次
相关 Paper
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing 等CVPR 2023
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
- Progressively Compressed Auto-Encoder for Self-supervised Representation LearningJin Li, Yaoming Wang, Xiaopeng Zhang, Yabo Chen 等ICLR 2023
- Task-customized Masked Autoencoder via Mixture of Cluster-conditional ExpertsZhili Liu, Kai Chen, Jianhua Han, Lanqing Hong 等ICLR 2023 · 被引用 6 次
- The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-trainingHao Liu, Xinghua Jiang, Xin Li, Antai Guo 等AAAI 2023 · 被引用 45 次
