Domain Knowledge Enhanced Vision-Language Pretrained Model for Dynamic Facial Expression Recognition
Liupeng Li, Yuhua Zheng, Shupeng Liu, Xiaoyin Xu, Taihao Li
Abstract
Dynamic facial expression recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video sequences. However, the complex temporal modeling caused by noisy frames, along with the limited training data significantly hinder the further development of DFER. Previous efforts in this domain have been limited as they tackled these issues separately. Inspired by recent advances of pretrained vision-language models (e.g., CLIP), we propose to leverage it to jointly address the two limitations in DFER. Since the raw CLIP model lacks the ability to model temporal relationships and determine the optimal task-related textual prompts, we utilize DFER-specific domain knowledge, including characteristics of temporal correlations and relationships between facial behavior descriptions at different levels, to guide the adaptation of CLIP to DFER. Specifically, we propose enhancements to CLIP's visual encoder through the design of a hierarchical video encoder that captures both short- and long-term temporal correlations in DFER. Meanwhile, we align facial expressions with action units through prior knowledge to construct semantically rich textual prompts, which are further enhanced with visual contents. Furthermore, we introduce a class-aware consistency regularization mechanism that adaptively filters out noisy frames, bolstering the model's robustness against interference. Extensive experiments on three in-the-wild dynamic facial expression datasets demonstrate that our method outperforms the state-of-the-art DFER approaches. The code is available at https://github.com/liliupeng28/DK-CLIP.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c2e54125-9f4c-46a8-aea3-8776c1978b70Related papers
- FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERsHaodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng et al.ACM MM 2024 · 26 citations
- Multimodal Prompt Alignment for Facial Expression RecognitionFuyan Ma, Yiran He, Bin Sun, Shutao LiICCV 2025 · 5 citations
- Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingYuanyuan Liu, Yuxuan Huang, Shuyang Liu, Yibing Zhan et al.ACM MM 2024 · 15 citations
- From Pixels to Semantics: Unified Facial Action Representation Learning for Micro-Expression AnalysisYicheng Deng, Hideaki Hayashi, Hajime NagaharaICLR 2026
- Enhanced Motion-Text Alignment for Image-to-Video Transfer LearningWei Zhang, Chaoqun Wan, Tongliang Liu, Xinmei Tian et al.CVPR 2024 · 8 citations
