Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior Understanding
Xiang Zhang, Taoyue Wang, Xiaotian Li, Huiyuan Yang, Lijun Yin
Abstract
Contrastive learning has shown promising potential for learning robust representations by utilizing unlabeled data. However, constructing effective positive-negative pairs for contrastive learning on facial behavior datasets remains challenging. This is because such pairs inevitably encode the subject-ID information, and the randomly constructed pairs may push similar facial images away due to the limited number of subjects in facial behavior datasets. To address this issue, we propose to utilize activity descriptions, coarse-grained information provided in some datasets, which can provide high-level semantic information about the image sequences but is often neglected in previous studies. More specifically, we introduce a two-stage Contrastive Learning with Text-Embeded framework for Facial behavior understanding (CLEF). The first stage is a weakly-supervised contrastive learning method that learns representations from positive-negative pairs constructed using coarse-grained activity information. The second stage aims to train the recognition of facial expressions or facial action units by maximizing the similarity between the image and the corresponding text label names. The proposed CLEF achieves state-of-the-art performance on three in-the-lab datasets for AU recognition and three in-the-wild datasets for facial expression recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bee133b-3bd5-4b9a-b324-886626a8358cCited by top-tier papers4
- Multimodal Prompt Alignment for Facial Expression RecognitionFuyan Ma, Yiran He, Bin Sun, Shutao LiICCV 2025 · 5 citations
- AURA: Visually Interpretable Affective Understanding via Robust ArchetypesGuanyu Hu, Dimitrios Kollias, Xinyu YangICML 2026
- Self-Supervised Facial Representation Learning with Facial Region AwarenessZheng Gao, Ioannis PatrasCVPR 2024
- MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationXiangdong Li, Ye Lou, Ao Gao, Wei Zhang et al.AAAI 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- Knowledge-Driven Self-Supervised Representation Learning for Facial Action Unit RecognitionYanan Chang, Shangfei WangCVPR 2022 · 38 citations
- Pursuing Knowledge Consistency: Supervised Hierarchical Contrastive Learning for Facial Action Unit RecognitionYingjie Chen, Chong Chen, Xiao Luo, Jianqiang Huang et al.ACM MM 2022 · 5 citations
- Pose-disentangled Contrastive Learning for Self-supervised Facial RepresentationYuanyuan Liu, Wenbin Wang, Yibing Zhan, Shaoze Feng et al.CVPR 2023
- Multi-Label Compound Expression Recognition: C-EXPR Database & NetworkDimitrios KolliasCVPR 2023
- Knowledge Augmented Deep Neural Networks for Joint Facial Expression and Action Unit RecognitionZijun Cui, Tengfei Song, Yuru Wang, Qiang JiNeurIPS 2020 · 70 citations
