Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
Xuri Ge, Junchen Fu, Fuhai Chen, Shan An, Nicu Sebe, Joemon M. Jose
摘要
Facial action units (AUs), as defined in the Facial Action Coding System (FACS), have received significant research interest owing to their diverse range of applications in facial state analysis. Current mainstream FAU recognition models have a notable limitation, i.e., focusing only on the accuracy of AU recognition and overlooking explanations of corresponding AU states. In this paper, we propose an end-to-end Vision-Language joint learning network for explainable FAU recognition (termed VL-FAU), which aims to reinforce AU representation capability and language interpretability through the integration of joint multimodal tasks. Specifically, VL-FAU brings together language models to generate fine-grained local muscle descriptions and distinguishable global face description when optimising FAU recognition. Through this, the global facial representation and its local AU representations will achieve higher distinguishability among different AUs and different subjects. In addition, multi-level AU representation learning is utilised to improve AU individual attention-aware representation capabilities based on multi-scale combined facial stem feature. Extensive experiments on DISFA and BP4D AU datasets show that the proposed approach achieves superior performance over the state-of-the-art methods on most of the metrics. In addition, compared with mainstream FAU recognition methods, VL-FAU can provide local- and global-level interpretability language descriptions with the AUs' predictions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li 等EMNLP 2021 · 被引用 130 次
- Uncertain Graph Neural Networks for Facial Action Unit DetectionTengfei Song, Lisha Chen, Wenming Zheng, Qiang JiAAAI 2021 · 被引用 86 次
- Knowledge Augmented Deep Neural Networks for Joint Facial Expression and Action Unit RecognitionZijun Cui, Tengfei Song, Yuru Wang, Qiang JiNeurIPS 2020 · 被引用 70 次
- Structured Multi-modal Feature Embedding and Alignment for Image-Sentence RetrievalXuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji 等ACM MM 2021 · 被引用 49 次
相关 Paper
- FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and ReasoningZhuozhao Hu, Kaishen Yuan, Xin Liu, Zitong Yu 等ACM MM 2025 · 被引用 3 次
- Facial Action Unit Detection With TransformersGeethu Miriam Jacob, Björn StengerCVPR 2021
- Knowledge-Driven Self-Supervised Representation Learning for Facial Action Unit RecognitionYanan Chang, Shangfei WangCVPR 2022 · 被引用 38 次
- CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit RecognitionYingjie Chen, Diqi Chen, Yizhou Wang, Tao Wang 等ACM MM 2021 · 被引用 10 次
- Facial-R1: Aligning Reasoning and Recognition for Facial Emotion AnalysisJiulong Wu, Yucheng Shen, Lingyong Yan, Haixin Sun 等AAAI 2026 · 被引用 3 次
