R-MAE: Regions Meet Masked Autoencoders
Duy-Kien Nguyen, Yanghao Li, Vaibhav Aggarwal, Martin R. Oswald, Alexander Kirillov, Cees G. M. Snoek, Xinlei Chen
Abstract
In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8eea25e2-6d18-4c93-af91-23cd5e666cc6Cited by top-tier papers6
- OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang et al.NeurIPS 2024 · 45 citations
- Feed-Forward SceneDINO for Unsupervised Semantic Scene CompletionAleksandar Jevtic, Christoph Reich, Felix Wimbauer, Oliver Hahn et al.ICCV 2025 · 3 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
- FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation LearningGaojian Wang, Feng Lin, Tong Wu, Zhenguang Liu et al.CVPR 2025
- Scene-Centric Unsupervised Panoptic SegmentationOliver Hahn, Christoph Reich, Nikita Araslanov, Daniel Cremers et al.CVPR 2025
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- SemMAE: Semantic-Guided Masking for Learning Masked AutoencodersGang Li, Heliang Zheng, Daqing Liu, Chaoyue Wang et al.NeurIPS 2022 · 188 citations
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
