DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing
Xiaoqin Wang, Xianxu Hou, Meidan Ding, Junliang Chen, Kaijun Deng, Jinheng Xie, Linlin Shen
Abstract
Face parsing aims to segment facial images into key components such as eyes, lips, and eyebrows. While existing methods rely on dense pixel-level annotations, such annotations are expensive and labor-intensive to obtain. To reduce annotation cost, we introduce Weakly Supervised Face Parsing (WSFP), a new task setting that performs dense facial component segmentation using only weak supervision, such as image-level labels and natural language descriptions. WSFP introduces unique challenges due to the high co-occurrence and visual similarity of facial components, which lead to ambiguous activations and degraded parsing performance. To address this, we propose DisFaceRep, a representation disentanglement framework designed to separate co-occurring facial components through both explicit and implicit mechanisms. Specifically, we introduce a co-occurring component disentanglement strategy to explicitly reduce dataset-level bias, and a text-guided component disentanglement loss to guide component separation using language supervision implicitly. Extensive experiments on CelebAMask-HQ, LaPa, and Helen demonstrate the difficulty of WSFP and the effectiveness of DisFaceRep, which significantly outperforms existing weakly supervised semantic segmentation methods. The code will be released at https://github.com/CVI-SZU/DisFaceRep.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- FineXtrol: Controllable Motion Generation via Fine-Grained TextKeming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang et al.AAAI 2026 · 3 citations
- SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image SegmentationMeihua Li, Yang Zhang, Weizhao He, Hu Qu et al.CVPR 2026 · 1 citation
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu et al.ICLR 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 710 citations
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.CVPR 2022 · 275 citations
- Integral Object Mining via Online Attention AccumulationPeng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng et al.ICCV 2019 · 246 citations
Related papers
- Dual-Structure Disentangling Variational Generation for Data-Limited Face ParsingPeipei Li, Yinglu Liu, Hailin Shi, Xiang Wu et al.ACM MM 2020 · 8 citations
- A New Dataset and Boundary-Attention Semantic Segmentation for Face ParsingYinglu Liu, Hailin Shi, Hao Shen, Yue Si et al.AAAI 2020 · 88 citations
- Parameter Efficient Local Implicit Image Function Network for Face SegmentationMausoom Sarkar, Nikitha S. R., Mayur Hemani, Rishabh Jain et al.CVPR 2023
- Decoupled Multi-task Learning with Cyclical Self-Regulation for Face ParsingQingping Zheng, Jiankang Deng, Zheng Zhu, Ying Li et al.CVPR 2022 · 46 citations
- Unsupervised Disentanglement of Linear-Encoded Facial SemanticsYutong Zheng, Yu-Kai Huang, Ran Tao, Zhiqiang Shen et al.CVPR 2021
