Toward High Quality Facial Representation Learning
Yue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi, Liang Liu, Yabiao Wang, Chengjie Wang
Abstract
Face analysis tasks have a wide range of applications, but the universal facial representation has only been explored in a few works. In this paper, we explore high-performance pre-training methods to boost the face analysis tasks such as face alignment and face parsing. We propose a self-supervised pre-training framework, called Mask Contrastive Face (MCF), with mask image modeling and a contrastive strategy specially adjusted for face domain tasks. To improve the facial representation quality, we use feature map of a pre-trained visual backbone as a supervision item and use a partially pre-trained decoder for mask image modeling. To handle the face identity during the pre-training stage, we further use random masks to build contrastive learning pairs. We conduct the pre-training on the LAION-FACE-cropped dataset, a variants of LAION-FACE 20M, which contains more than 20 million face images from Internet websites. For efficiency pre-training, we explore our framework pre-training performance on a small part of LAION-FACE-cropped and verify the superiority with different pre-training settings. Our model pre-trained with the full pre-training dataset outperforms the state-of-the-art methods on multiple downstream tasks. Our model achieves 0.932 NME_diag for AFLW-19 face alignment and 93.96 F1 score for LaPa face parsing. Code is available at https://github.com/nomewang/MCF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8307e92d-0e5a-4a84-ba3a-3dac13d96483Cited by top-tier papers4
- SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataXilin He, Cheng Luo, Xiaole Xian, Bing Li et al.ICCV 2025 · 6 citations
- FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation LearningGaojian Wang, Feng Lin, Tong Wu, Zhenguang Liu et al.CVPR 2025
- Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition DatasetsZhichao Chen, Yongle Zhao, Kaicheng Yang, Meng Yang et al.ICML 2026
- Self-Supervised Facial Representation Learning with Facial Region AwarenessZheng Gao, Ioannis PatrasCVPR 2024
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- General Facial Representation Learning in a Visual-Linguistic MannerYinglin Zheng, Hao Yang, Ting Zhang, Jianmin Bao et al.CVPR 2022 · 161 citations
- PrefAce: Face-Centric Pretraining with Self-Structure Aware DistillationSiyuan Hu, Zheng Wang, Peng Hu, Xi Peng et al.AAAI 2024 · 2 citations
- A New Dataset and Boundary-Attention Semantic Segmentation for Face ParsingYinglu Liu, Hailin Shi, Hao Shen, Yue Si et al.AAAI 2020 · 88 citations
- FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-TrainingYudong Li, Xianxu Hou, Dezhi Zheng, Linlin Shen et al.ACM MM 2024 · 2 citations
- LAFS: Landmark-Based Facial Self-Supervised Learning for Face RecognitionZhonglin Sun, Chen Feng, Ioannis Patras, Georgios TzimiropoulosCVPR 2024 · 17 citations
