TransFER: Learning Relation-aware Facial Expression Representations with Transformers
Fanglei Xue, Qiangchang Wang, Guodong Guo
Abstract
Facial expression recognition (FER) has received increasing interest in computer vision. We propose the Trans-FER model which can learn rich relation-aware local representations. It mainly consists of three components: Multi-Attention Dropping (MAD), ViT-FER, and Multi-head Self-Attention Dropping (MSAD). First, local patches play an important role in distinguishing various expressions, however, few existing works can locate discriminative and diverse local patches. This can cause serious problems when some patches are invisible due to pose variations or viewpoint changes. To address this issue, the MAD is proposed to randomly drop an attention map. Consequently, models are pushed to explore diverse local patches adaptively. Second, to build rich relations between different local patches, the Vision Transformers (ViT) are used in FER, called ViT-FER. Since the global scope is used to reinforce each local patch, a better representation is obtained to boost the FER performance. Thirdly, the multi-head self-attention allows ViT to jointly attend to features from different information subspaces at different positions. Given no explicit guidance, however, multiple self-attentions may extract similar relations. To address this, the MSAD is proposed to randomly drop one self-attention module. As a result, models are forced to learn rich relations among diverse local patches. Our proposed TransFER model outperforms the state-of-the-art methods on several FER benchmarks, showing its effectiveness and usefulness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28da04e9-fbcb-4777-a633-fa1784f3f705Cited by top-tier papers20
- Face2Exp: Combating Data Biases for Facial Expression RecognitionDan Zeng, Zhiyuan Lin, Xiao Yan, Yuting Liu et al.CVPR 2022 · 125 citations
- Towards Semi-Supervised Deep Facial Expression Recognition with An Adaptive Confidence MarginHangyu Li, Nannan Wang, Xi Yang, Xiaoyu Wang et al.CVPR 2022 · 97 citations
- Intensity-Aware Loss for Dynamic Facial Expression Recognition in the WildHanting Li, Hongjing Niu, Zhaoqing Zhu, Feng ZhaoAAAI 2023 · 94 citations
- LA-Net: Landmark-Aware Learning for Reliable Facial Expression Recognition under Label NoiseZhiyu Wu, Jinshi CuiICCV 2023 · 47 citations
- Leave No Stone Unturned: Mine Extra Knowledge for Imbalanced Facial Expression RecognitionYuhang Zhang, Yaqi Li, Lixiong Qin, Xuannan Liu et al.NeurIPS 2023 · 47 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Label Distribution Learning on Auxiliary Label Space Graphs for Facial Expression RecognitionShikai Chen, Jianfeng Wang, Yuedong Chen, Zhongchao Shi et al.CVPR 2020
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
- Suppressing Uncertainties for Large-Scale Facial Expression RecognitionKai Wang, Xiaojiang Peng, Jianfei Yang, Shijian Lu et al.CVPR 2020
Related papers
- MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression RecognitionLicai Sun, Zheng Lian, Bin Liu, Jianhua TaoACM MM 2023 · 85 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Latent-OFER: Detect, Mask, and Reconstruct with Latent Vectors for Occluded Facial Expression RecognitionIsack Lee, Eungi Lee, Seok Bong YooICCV 2023 · 41 citations
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression RecognitionYu Gao, Haoyu Ji, Zhiyong Wang, Wenze Huang et al.AAAI 2026
