Adaptive Multimodal Fusion for Facial Action Units Recognition
Huiyuan Yang, Taoyue Wang, Lijun Yin
Abstract
Multimodal facial action units (AU) recognition aims to build models that are capable of processing, correlating, and integrating information from multiple modalities (i.e., 2D images from a visual sensor, 3D geometry from 3D imaging, and thermal images from an infrared sensor). Although the multimodel data can provide rich information, there are two challenges that have to be addressed when learning from multimodal data: 1) the model must capture the complex cross-modal interactions in order to utilize the additional and mutual information effectively; 2) the model must be robust enough in the circumstance of unexpected data corruptions during testing, in case of a certain modality missing or being noisy. In this paper, we propose a novel A daptive M ultimodal F usion method (AMF ) for AU detection, which learns to select the most relevant feature representations from different modalities by a re-sampling procedure conditioned on a feature scoring module. The feature scoring module is designed to allow for evaluating the quality of features learned from multiple modalities. As a result, AMF is able to adaptively select more discriminative features, thus increasing the robustness to missing or corrupted modalities. In addition, to alleviate the over-fitting problem and make the model generalize better on the testing data, a cut-switch multimodal data augmentation method is designed, by which a random block is cut and switched across multiple modalities. We have conducted a thorough investigation on two public multimodal AU datasets, BP4D and BP4D+, and the results demonstrate the effectiveness of the proposed method. Ablation studies on various circumstances also show that our method remains robust to missing or noisy modalities during tests.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior UnderstandingXiang Zhang, Taoyue Wang, Xiaotian Li, Huiyuan Yang et al.ICCV 2023 · 26 citations
- Knowledge Perceived Multi-modal Pretraining in E-commerceYushan Zhu, Huaixiao Zhao, Wen Zhang, Ganqiang Ye et al.ACM MM 2021 · 21 citations
- ReactioNet: Learning High-order Facial Behavior from Universal Stimulus-Reaction by Dyadic Relation ReasoningXiaotian Li, Taoyue Wang, Geran Zhao, Xiang Zhang et al.ICCV 2023 · 3 citations
- Biomechanics-Guided Facial Action Unit Detection Through Force ModelingZijun Cui, Chenyi Kuang, Tian Gao, Kartik Talamadupula et al.CVPR 2023
Related papers
- PIAP-DF: Pixel-Interested and Anti Person-Specific Facial Action Unit Detection Net with Discrete Feedback LearningYang Tang, Wangding Zeng, Dafei Zhao, Honggang ZhangICCV 2021 · 39 citations
- CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit RecognitionYingjie Chen, Diqi Chen, Yizhou Wang, Tao Wang et al.ACM MM 2021 · 10 citations
- Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing ModalitiesYan Zhuang, Minhao Liu, Wei Bai, Yanru Zhang et al.NeurIPS 2025 · 10 citations
- Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesChen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi et al.ICCV 2025 · 8 citations
- Self-Supervised Regional and Temporal Auxiliary Tasks for Facial Action Unit RecognitionJingwei Yan, Jingjing Wang, Qiang Li, Chunmao Wang et al.ACM MM 2021 · 9 citations
