SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark Estimation
Kejia Yin, Varshanth S. Rao, Ruowei Jiang, Xudong Liu, Parham Aarabi, David B. Lindell
Abstract
Self-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task, existing state-of-the-art (SOTA) methods (1) extract coarse features from backbones that are trained with instance-level self-supervised learning (SSL) paradigms, which neglect the dense prediction nature of the task, (2) aggregate them into memory-intensive hypercolumn formations, and (3) su-pervise lightweight projector networks to naï vely establish full local correspondences among all pairs of spatial features. In this paper, we introduce SCE-MAE, a framework that (1) leverages the MAE [14], a region-level SSL method that naturally better suits the landmark prediction task, (2) operates on the vanilla feature map instead of on expen-sive hypercolumns, and (3) employs a Correspondence Ap-proximation and Refinement Block (CARB) that utilizes a simple density peak clustering algorithm and our proposed Locality-Constrained Repellence Loss to directly hone only select local correspondences. We demonstrate through extensive experiments that SCE-MAE is highly effective and robust, outperforming existing SOTA methods by large mar-gins of 20%-44% on the landmark matching and 9%-15% on the landmark detection tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb4c2add-1a1a-4f0c-ac11-795c919fcdd4Cited by top-tier papers1
Ask how each one uses itBuilds on25
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Exploiting Self-Supervised and Semi-Supervised Learning for Facial Landmark Tracking with Unlabeled DataShi Yin, Shangfei Wang, Xiaoping Chen, Enhong ChenACM MM 2020 · 7 citations
- Learning Transformation-Predictive Representations for Detection and Description of Local FeaturesZihao Wang, Chunxu Wu, Yifei Yang, Zhen LiCVPR 2023
- Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingXinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong et al.CVPR 2021
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised LearningXuecheng Wu, Heli Sun, Yifan Wang, Jiayu Nie et al.CVPR 2025
