QCS: Feature Refining from Quadruplet Cross Similarity for Facial Expression Recognition
Chengpeng Wang, Li Chen, Lili Wang, Zhaofan Li, Xuebin Lv
Abstract
Facial expression recognition faces challenges where labeled significant features in datasets are mixed with unlabeled redundant ones. In this paper, we introduce Cross Similarity Attention (CSA) to mine richer intrinsic information from image pairs, overcoming a limitation when the Scaled Dot-Product Attention of ViT is directly applied to calculate the similarity between two different images. Based on CSA, we simultaneously minimize intra-class differences and maximize inter-class differences at the fine-grained feature level through interactions among multiple branches. Contrastive residual distillation is utilized to transfer the information learned in the cross module back to the base network. We ingeniously design a four-branch centrally symmetric network, named Quadruplet Cross Similarity (QCS), which alleviates gradient conflicts arising from the cross module and achieves balanced and stable training. It can adaptively extract discriminative features while isolating redundant ones. The cross-attention modules exist during training, and only one base branch is retained during inference, resulting in no increase in inference time. Extensive experiments show that our proposed method achieves state-of-the-art performance on several FER datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Facial-R1: Aligning Reasoning and Recognition for Facial Emotion AnalysisJiulong Wu, Yucheng Shen, Lingyong Yan, Haixin Sun et al.AAAI 2026 · 3 citations
- ContextFace: Generating Facial Expressions from Emotional ContextsMin-Jung Kim, Minsang Kim, Seung Jun BaekICCV 2025 · 1 citation
- HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression RecognitionYu Gao, Haoyu Ji, Zhiyong Wang, Wenze Huang et al.AAAI 2026
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
Related papers
- TransFER: Learning Relation-aware Facial Expression Representations with TransformersFanglei Xue, Qiangchang Wang, Guodong GuoICCV 2021 · 276 citations
- When Facial Expression Recognition Meets Few-Shot Learning: A Joint and Alternate Learning FrameworkXinyi Zou, Yan Yan, Jing-Hao Xue, Si Chen et al.AAAI 2022 · 22 citations
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu et al.CVPR 2022 · 251 citations
- Facial Action Unit Detection With TransformersGeethu Miriam Jacob, Björn StengerCVPR 2021
- Learning with Alignments: Tackling the Inter- and Intra-domain Shifts for Cross-multidomain Facial Expression RecognitionYuxiang Yang, Lu Wen, Xinyi Zeng, Yuanyuan Xu et al.ACM MM 2024 · 7 citations
