Patch-Aware Representation Learning for Facial Expression Recognition
Yi Wu, Shangfei Wang, Yanan Chang
Abstract
Existing methods for facial expression recognition (FER) lack the utilization of prior facial knowledge, primarily focusing on expression-related regions while disregarding explicitly processing expression-independent information. This paper proposes a patch-aware FER method that incorporates facial keypoints to guide the model and learns precise representations through two collaborative streams, addressing these issues. First, facial keypoints are detected using a facial landmark detection algorithm, and the facial image is divided into equal-sized patches using the Patch Embedding Module. Then, a correlation is established between the keypoints and patches using a simplified conversion relationship. Two collaborative streams are introduced, each corresponding to a specific mask strategy. The first stream masks patches corresponding to the keypoints, excluding those along the facial contour, with a certain probability. The resulting image embedding is input into the Encoder to obtain expression-related features. The features are passed through the Decoder and Classifier to reconstruct the masked patches and recognize the expression, respectively. The second stream masks patches corresponding to all the above keypoints. The resulting image embedding is input into the Encoder and Classifier successively, with the resulting logit approximating a uniform distribution. Through the first stream, the Encoder learns features in the regions related to expression, while the second stream enables the Encoder to better ignore expression-independent information, such as the background, facial contours, and hair. Experiments on two benchmark datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 56dbe779-ab38-4cc7-b7c8-fff18d377a1fCited by top-tier papers4
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang et al.ICCV 2025 · 21 citations
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationFeng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang et al.ACM MM 2025 · 10 citations
- Learning with Alignments: Tackling the Inter- and Intra-domain Shifts for Cross-multidomain Facial Expression RecognitionYuxiang Yang, Lu Wen, Xinyi Zeng, Yuanyuan Xu et al.ACM MM 2024 · 7 citations
- Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context TransformerXinpeng Li, Teng Wang, Jian Zhao, Shuyi Mao et al.ACM MM 2024 · 3 citations
Related papers
- DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosYan Wang, Yixuan Sun, Wei Song, Shuyong Gao et al.ACM MM 2022 · 62 citations
- Knowledge Augmented Deep Neural Networks for Joint Facial Expression and Action Unit RecognitionZijun Cui, Tengfei Song, Yuru Wang, Qiang JiNeurIPS 2020 · 70 citations
- Context-Aware Feature and Label Fusion for Facial Action Unit Intensity Estimation With Partially Labeled DataYong Zhang, Haiyong Jiang, Baoyuan Wu, Yanbo Fan et al.ICCV 2019 · 32 citations
- Latent-OFER: Detect, Mask, and Reconstruct with Latent Vectors for Occluded Facial Expression RecognitionIsack Lee, Eungi Lee, Seok Bong YooICCV 2023 · 41 citations
- Uncertainty-aware Cross-dataset Facial Expression Recognition via Regularized Conditional AlignmentLinyi Zhou, Xijian Fan, Yingjie Ma, Tardi Tjahjadi et al.ACM MM 2020 · 20 citations
