Understanding Masked Autoencoders via Hierarchical Latent Variable Models
Lingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing, Yuejie Chi, Louis-Philippe Morency, Kun Zhang
摘要
Masked autoencoder (MAE), a simple and effective selfsupervised learning framework based on the reconstruction of masked image regions, has recently achieved prominent success in a variety of vision tasks. Despite the emergence of intriguing empirical observations on MAE, a theoretically principled understanding is still lacking. In this work, we formally characterize and justify existing empirical insights and provide theoretical guarantees of MAE. We formulate the underlying data-generating process as a hierarchical latent variable model, and show that under reasonable assumptions, MAE provably identifies a set of latent variables in the hierarchical model, explaining why MAE can extract high-level information from pixels. Further, we show how key hyperparameters in MAE (the masking ratio and the patch size) determine which true latent variables to be recovered, therefore influencing the level of semantic information in the representation. Specifically, extremely large or small masking ratios inevitably lead to low-level representations. Our theory offers coherent explanations of existing empirical observations and provides insights for potential empirical improvements and fundamental limitations of the masked-reconstruction paradigm. We conduct extensive experiments to validate our theoretical insights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time SeriesSimon A. Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee 等ICLR 2026 · 被引用 22 次
- Information Flow in Self-Supervised LearningZhiquan Tan, Jingqin Yang, Weiran Huang, Yang Yuan 等ICML 2024 · 被引用 18 次
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 被引用 18 次
- Counterfactual Generation with Identifiability GuaranteesHanqi Yan, Lingjing Kong, Lin Gui, Yuejie Chi 等NeurIPS 2023 · 被引用 16 次
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 被引用 12 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial CorrelationsAnthony Bisulco, Rahul Ramesh, Randall Balestriero, Pratik ChaudhariICCV 2025 · 被引用 2 次
- Learning Mask Invariant Mutual Information for Masked Image ModelingTao Huang, Yanxiang Ma, Shan You, Chang XuICLR 2025
- How Mask Matters: Towards Theoretical Understandings of Masked AutoencodersQi Zhang, Yifei Wang, Yisen WangNeurIPS 2022 · 被引用 119 次
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
- Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized ConstraintsJiahao Xia, Yike Wu, Wenjian Huang, Jianguo Zhang 等ICCV 2025 · 被引用 1 次
