SegFace: Face Segmentation of Long-Tail Classes
Kartik Narayan, Vibashan VS, Vishal M. Patel
Abstract
Face parsing refers to the semantic segmentation of human faces into key facial regions such as eyes, nose, hair, etc. It serves as a prerequisite for various advanced applications, including face editing, face swapping, and facial makeup, which often require segmentation masks for classes like eyeglasses, hats, earrings, and necklaces. These infrequently occurring classes are called long-tail classes, which are overshadowed by more frequently occurring classes known as head classes. Existing methods, primarily CNN-based, tend to be dominated by head classes during training, resulting in suboptimal representation for long-tail classes. Previous works have largely overlooked the problem of poor segmentation performance of long-tail classes. To address this issue, we propose SegFace, a simple and efficient approach that uses a lightweight transformer-based model which utilizes learnable class-specific tokens. The transformer decoder leverages class-specific tokens, allowing each token to focus on its corresponding class, thereby enabling independent modeling of each class. The proposed approach improves the performance of long-tail classes, thereby boosting overall performance. To the best of our knowledge, SegFace is the first work to employ transformer models for face parsing. Moreover, our approach can be adapted for low-compute edge devices, achieving 95.96 FPS. We conduct extensive experiments demonstrating that SegFace significantly outperforms previous state-of-the-art models, achieving a mean F1 score of 88.96 (+2.82) on the CelebAMask-HQ dataset and 93.03 (+0.65) on the LaPa dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 446d4648-a8ea-4934-93f5-c15fd1103338Cited by top-tier papers7
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsDingqiang Ye, Chao Fan, Zhanbo Huang, Chengwen Luo et al.NeurIPS 2025 · 28 citations
- FaceXFormer: A Unified Transformer for Facial AnalysisKartik Narayan, Vibashan VS, Rama Chellappa, Vishal M. PatelICCV 2025 · 16 citations
- Dynamic Dictionary Learning for Remote Sensing Image SegmentationXuechao Zou, Yue Li, Shun Zhang, Kai Li et al.ICCV 2025 · 15 citations
- Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go BeyondHuiyu Zhai, Xingxing Yang, Yalan Ye, Chenyang Li et al.ACM MM 2025 · 5 citations
- Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative ModelsFrancisco Caetano, Christiaan G. A. Viviers, Peter H. N. de With, Fons van der SommenAAAI 2026 · 4 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- General Facial Representation Learning in a Visual-Linguistic MannerYinglin Zheng, Hao Yang, Ting Zhang, Jianmin Bao et al.CVPR 2022 · 161 citations
- A New Dataset and Boundary-Attention Semantic Segmentation for Face ParsingYinglu Liu, Hailin Shi, Hao Shen, Yue Si et al.AAAI 2020 · 88 citations
Related papers
- Parameter Efficient Local Implicit Image Function Network for Face SegmentationMausoom Sarkar, Nikitha S. R., Mayur Hemani, Rishabh Jain et al.CVPR 2023
- SOTR: Segmenting Objects with TransformersRuohao Guo, Dantong Niu, Liao Qu, Zhenbo LiICCV 2021 · 123 citations
- PEM: Prototype-Based Efficient MaskFormer for Image SegmentationNiccolò Cavagnero, Gabriele Rosi, Claudia Cuttano, Francesca Pistilli et al.CVPR 2024
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time SegmentationZhengze Xu, Dongyue Wu, Changqian Yu, Xiangxiang Chu et al.AAAI 2024 · 166 citations
