Self-Supervised Facial Representation Learning with Facial Region Awareness
Zheng Gao, Ioannis Patras
Abstract
Self-supervised pre-training has been proved to be effective in learning transferable representations that benefit various visual tasks. This paper asks this question: can self-supervised pre-training learn general facial representations for various facial analysis tasks? Recent efforts toward this goal are limited to treating each face image as a whole, i.e., learning consistent facial representations at the image-level, which overlooks the "consistency of local facial representations" (i.e., facial regions like eyes, nose, etc). In this work, we make a first attempt to propose a novel self-supervised facial representation learning framework to learn consistent global and local facial representations, Facial Region Awareness (FRA). Specifically, we explicitly enforce the consistency of facial regions by matching the local facial representations across views, which are extracted with learned heatmaps highlighting the facial regions. Inspired by the mask prediction in supervised semantic segmentation, we obtain the heatmaps via cosine similarity between the per-pixel projection of feature maps and "facial mask embeddings" computed from learnable positional embeddings, which leverage the attention mechanism to globally look up the facial image for facial regions. To learn such heatmaps, we formulate the learning of facial mask embeddings as a deep clustering problem by assigning the pixel features from the feature maps to them. The transfer learning results on facial classification and regression tasks show that our FRA outperforms previous pre-trained models and more importantly, using ResNet as the unified backbone for various tasks, our FRA achieves comparable or even better performance compared with SOTA methods in facial analysis tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9ccae7f-acad-4a94-a4a7-3adad7bf6446Cited by top-tier papers7
- CLIPCleaner: Cleaning Noisy Labels with CLIPChen Feng, Georgios Tzimiropoulos, Ioannis PatrasACM MM 2024 · 12 citations
- QCS: Feature Refining from Quadruplet Cross Similarity for Facial Expression RecognitionChengpeng Wang, Li Chen, Lili Wang, Zhaofan Li et al.AAAI 2025 · 9 citations
- SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataXilin He, Cheng Luo, Xiaole Xian, Bing Li et al.ICCV 2025 · 6 citations
- Heatmap Regression without Soft-Argmax for Facial Landmark DetectionChiao-An Yang, Raymond A. YehICCV 2025 · 3 citations
- Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup FeaturesZheng Gao, Debin Meng, Yunqi Miao, Zhensong Zhang et al.CVPR 2026 · 1 citation
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- PrefAce: Face-Centric Pretraining with Self-Structure Aware DistillationSiyuan Hu, Zheng Wang, Peng Hu, Xi Peng et al.AAAI 2024 · 2 citations
- MARLIN: Masked Autoencoder for facial video Representation LearnINgZhixi Cai, Shreya Ghosh, Kalin Stefanov, Abhinav Dhall et al.CVPR 2023
- Exploiting Self-Supervised and Semi-Supervised Learning for Facial Landmark Tracking with Unlabeled DataShi Yin, Shangfei Wang, Xiaoping Chen, Enhong ChenACM MM 2020 · 7 citations
- General Facial Representation Learning in a Visual-Linguistic MannerYinglin Zheng, Hao Yang, Ting Zhang, Jianmin Bao et al.CVPR 2022 · 161 citations
- Toward High Quality Facial Representation LearningYue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi et al.ACM MM 2023 · 4 citations
