Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go Beyond
Huiyu Zhai, Xingxing Yang, Yalan Ye, Chenyang Li, Bin Fan, Changze Li
摘要
Facial expression recognition (FER) is a challenging task due to pervasive occlusion and dataset biases. Especially when facial information is partially occluded, existing FER models struggle to extract effective facial features, leading to inaccurate classifications. In response, we present ORSANet, which introduces the following three key contributions: First, we introduce auxiliary multi-modal semantic guidance to disambiguate facial occlusion and learn high-level semantic knowledge, which is two-fold: 1) we introduce semantic segmentation maps as dense semantics prior to generate semantics-enhanced facial representations; 2) we introduce facial landmarks as sparse geometric prior to mitigate intrinsic noises in FER, such as identity and gender biases. Second, to facilitate the effective incorporation of these two multi-modal priors, we customize a Multi-scale Cross-interaction Module (MCM) to adaptively fuse the landmark feature and semantics-enhanced representations within different scales. Third, we design a Dynamic Adversarial Repulsion Enhancement Loss (DARELoss) that dynamically adjusts the margins of ambiguous classes, further enhancing the model's ability to distinguish similar expressions. We further construct the first occlusion-oriented FER dataset to facilitate specialized robustness analysis on various real-world occlusion conditions, dubbed Occlu-FER. Extensive experiments on both public benchmarks and Occlu-FER demonstrate that our proposed ORSANet achieves SOTA recognition performance. Code is publicly available at https://github.com/Wenyuzhy/ORSANet-master.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu 等NeurIPS 2021 · 被引用 798 次
- Robust Lightweight Facial Expression Recognition Network with Label Distribution TrainingZengqun Zhao, Qingshan Liu, Feng ZhouAAAI 2021 · 被引用 300 次
- TransFER: Learning Relation-aware Facial Expression Representations with TransformersFanglei Xue, Qiangchang Wang, Guodong GuoICCV 2021 · 被引用 276 次
相关 Paper
- LA-Net: Landmark-Aware Learning for Reliable Facial Expression Recognition under Label NoiseZhiyu Wu, Jinshi CuiICCV 2023 · 被引用 47 次
- JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression RecognitionYingjian Li, Yingnan Gao, Bingzhi Chen, Zheng Zhang 等ACM MM 2021 · 被引用 21 次
- Uncertainty-aware Cross-dataset Facial Expression Recognition via Regularized Conditional AlignmentLinyi Zhou, Xijian Fan, Yingjie Ma, Tardi Tjahjadi 等ACM MM 2020 · 被引用 20 次
- Learning from More: Combating Uncertainty Cross-multidomain for Facial Expression RecognitionHanwei Liu, Huiling Cai, Qingcheng Lin, Xuefeng Li 等ACM MM 2023 · 被引用 5 次
- Occluded Facial Expression Recognition with Step-Wise Assistance from Unpaired Non-Occluded ImagesBin Xia, Shangfei WangACM MM 2020 · 被引用 20 次
