HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition
Yu Gao, Haoyu Ji, Zhiyong Wang, Wenze Huang, Qian Dong, Zhihao Yang, Xueting Liu, Weihong Ren, Honghai Liu
Abstract
Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different head poses, and so on. To address the above problems, current approaches rely on extensive learnable parameters and complex model architectures, which inevitably lead to overfitting and cause the FER model to focus on non-discriminative facial regions. In this work, we propose an HKAFER model that can adaptively enhance visual expression representations through efficiently fine-tuning the image encoder in large Visual Foundation Models (VFMs) and Vision-Language Models (VLMs). Specifically, we establish Heterogeneous Kronecker Adaptation (HeKA), which consists of multi-scale adapters based on Kronecker product in a parallel manner, offering significantly diverse subspaces to learn the incremental matrices. Besides, we also propose Dual-Branch Interactive Router (DBIR) to dynamically assign the weights of adapters, which promotes collaboration and information flow among them. In this way, our HKAFER can effectively capture robust spatial features and the regional associations. Experimental results demonstrate that our proposed model not only outperforms state-of-the-art methods on several FER benchmarks but also uses significantly fewer trainable parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f2aa412-0874-4745-9fe4-9c20ae8b7a73Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov et al.ICML 2024 · 820 citations
- TransFER: Learning Relation-aware Facial Expression Representations with TransformersFanglei Xue, Qiangchang Wang, Guodong GuoICCV 2021 · 276 citations
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam et al.AAAI 2023 · 234 citations
Related papers
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationFeng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang et al.ACM MM 2025 · 10 citations
- Robust Lightweight Facial Expression Recognition Network with Label Distribution TrainingZengqun Zhao, Qingshan Liu, Feng ZhouAAAI 2021 · 300 citations
- DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosYan Wang, Yixuan Sun, Wei Song, Shuyong Gao et al.ACM MM 2022 · 62 citations
- Variance-Aware Bi-Attention Expression Transformer for Open-Set Facial Expression Recognition in the WildJunjie Zhu, Bingjun Luo, Ao Sun, Jinghang Tan et al.ACM MM 2023 · 6 citations
- Latent-OFER: Detect, Mask, and Reconstruct with Latent Vectors for Occluded Facial Expression RecognitionIsack Lee, Eungi Lee, Seok Bong YooICCV 2023 · 41 citations
