HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition
Yu Gao, Haoyu Ji, Zhiyong Wang, Wenze Huang, Qian Dong, Zhihao Yang, Xueting Liu, Weihong Ren, Honghai Liu
摘要
Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different head poses, and so on. To address the above problems, current approaches rely on extensive learnable parameters and complex model architectures, which inevitably lead to overfitting and cause the FER model to focus on non-discriminative facial regions. In this work, we propose an HKAFER model that can adaptively enhance visual expression representations through efficiently fine-tuning the image encoder in large Visual Foundation Models (VFMs) and Vision-Language Models (VLMs). Specifically, we establish Heterogeneous Kronecker Adaptation (HeKA), which consists of multi-scale adapters based on Kronecker product in a parallel manner, offering significantly diverse subspaces to learn the incremental matrices. Besides, we also propose Dual-Branch Interactive Router (DBIR) to dynamically assign the weights of adapters, which promotes collaboration and information flow among them. In this way, our HKAFER can effectively capture robust spatial features and the regional associations. Experimental results demonstrate that our proposed model not only outperforms state-of-the-art methods on several FER benchmarks but also uses significantly fewer trainable parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 等ICML 2024 · 被引用 820 次
- TransFER: Learning Relation-aware Facial Expression Representations with TransformersFanglei Xue, Qiangchang Wang, Guodong GuoICCV 2021 · 被引用 276 次
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam 等AAAI 2023 · 被引用 234 次
相关 Paper
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationFeng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang 等ACM MM 2025 · 被引用 10 次
- Robust Lightweight Facial Expression Recognition Network with Label Distribution TrainingZengqun Zhao, Qingshan Liu, Feng ZhouAAAI 2021 · 被引用 300 次
- DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosYan Wang, Yixuan Sun, Wei Song, Shuyong Gao 等ACM MM 2022 · 被引用 62 次
- Variance-Aware Bi-Attention Expression Transformer for Open-Set Facial Expression Recognition in the WildJunjie Zhu, Bingjun Luo, Ao Sun, Jinghang Tan 等ACM MM 2023 · 被引用 6 次
- Latent-OFER: Detect, Mask, and Reconstruct with Latent Vectors for Occluded Facial Expression RecognitionIsack Lee, Eungi Lee, Seok Bong YooICCV 2023 · 被引用 41 次
