How Mask Matters: Towards Theoretical Understandings of Masked Autoencoders
Qi Zhang, Yifei Wang, Yisen Wang
Abstract
Masked Autoencoders (MAE) based on a reconstruction task have risen to be a promising paradigm for self-supervised learning (SSL) and achieve state-of-the-art performance across different benchmark datasets. However, despite its impressive empirical success, there is still limited theoretical understanding of it. In this paper, we propose a theoretical understanding of how masking matters for MAE to learn meaningful features. We establish a close connection between MAE and contrastive learning, which shows that MAE implicit aligns the mask-induced positive pairs. Built upon this connection, we develop the first downstream guarantees for MAE methods, and analyze the effect of mask ratio. Besides, as a result of the implicit alignment, we also point out the dimensional collapse issue of MAE, and propose a Uniformity-enhanced MAE (U-MAE) loss that can effectively address this issue and bring significant improvements on real-world datasets, including CIFAR-10, ImageNet-100, and ImageNet-1K. Code is available at (https://github.com/zhangq327/U-MAE).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9470b8db-6118-4ed1-9643-643fe8da6dc2Cited by top-tier papers48
- What's Behind the Mask: Understanding Masked Graph Modeling for Graph AutoencodersJintang Li, Ruofan Wu, Wangbin Sun, Liang Chen et al.KDD 2023 · 89 citations
- Simple and Asymmetric Graph Contrastive Learning without AugmentationsTeng Xiao, Huaisheng Zhu, Zhengyu Chen, Suhang WangNeurIPS 2023 · 86 citations
- Rethinking Graph Masked Autoencoders through Alignment and UniformityLiang Wang, Xiang Tao, Qiang Liu, Shu Wu et al.AAAI 2024 · 40 citations
- Representation Uncertainty in Self-Supervised Learning as Variational InferenceHiroki Nakamura, Masashi Okada, Tadahiro TaniguchiICCV 2023 · 27 citations
- Adversarial Examples Are Not Real FeaturesAng Li, Yifei Wang, Yiwen Guo, Yisen WangNeurIPS 2023 · 24 citations
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Learning Mask Invariant Mutual Information for Masked Image ModelingTao Huang, Yanxiang Ma, Shan You, Chang XuICLR 2025
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
- Generative and Contrastive Paradigms Are Complementary for Graph Self-Supervised LearningYuxiang Wang, Xiao Yan, Chuang Hu, Quanqing Xu et al.ICDE 2024 · 11 citations
- Modality-Agnostic Self-Supervised Learning with Meta-Learned Masked Auto-EncoderHuiwon Jang, Jihoon Tack, Daewon Choi, Jongheon Jeong et al.NeurIPS 2023 · 9 citations
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong et al.KDD 2022 · 533 citations
