Multi-Scale Dynamic and Hierarchical Relationship Modeling for Facial Action Units Recognition
Zihan Wang, Siyang Song, Cheng Luo, Songhe Deng, Weicheng Xie, Linlin Shen
Abstract
Human facial action units (AUs) are mutually related in a hierarchical manner, as not only they are associated with each other in both spatial and temporal domains but also AUs located in the same/close facial regions show stronger relationships than those of different facial regions. While none of existing approach thoroughly model such hierarchical inter-dependencies among AUs, this paper proposes to comprehensively model multi-scale AU-related dynamic and hierarchical spatio-temporal relationship among AUs for their occurrences recognition. Specifically, we first propose a novel multi-scale temporal differencing network with an adaptive weighting block to explicitly capture facial dynamics across frames at different spatial scales, which specifically considers the heterogeneity of range and magnitude in different AUs' activation. Then, a two-stage strategy is introduced to hierarchically model the relationship among AUs based on their spatial distribution (i.e., local and cross-region AU relationship modelling). Experimental results achieved on BP4D and DISFA show that our approach is the new state-of-the-art in the field of AU occurrence recognition. Our code is publicly available at https://github.com/CVI-SZU/MDHR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42efc9c8-ee48-4e42-96d0-ae5d53eda8b1Cited by top-tier papers4
- Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-SpeechRui Liu, Shuwei He, Yifan Hu, Haizhou LiAAAI 2025 · 8 citations
- AURA: Visually Interpretable Affective Understanding via Robust ArchetypesGuanyu Hu, Dimitrios Kollias, Xinyu YangICML 2026
- Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU DetectionYuru Wang, Yue ZhouCVPR 2026
- MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationXiangdong Li, Ye Lou, Ao Gao, Wei Zhang et al.AAAI 2026
Builds on10
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang et al.AAAI 2020 · 267 citations
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 189 citations
- Uncertain Graph Neural Networks for Facial Action Unit DetectionTengfei Song, Lisha Chen, Wenming Zheng, Qiang JiAAAI 2021 · 86 citations
- Knowledge Augmented Deep Neural Networks for Joint Facial Expression and Action Unit RecognitionZijun Cui, Tengfei Song, Yuru Wang, Qiang JiNeurIPS 2020 · 70 citations
- Knowledge-Driven Self-Supervised Representation Learning for Facial Action Unit RecognitionYanan Chang, Shangfei WangCVPR 2022 · 38 citations
Related papers
- Region of Interest Based Graph Convolution: A Heatmap Regression Approach for Action Unit DetectionZheng Zhang, Taoyue Wang, Lijun YinACM MM 2020 · 21 citations
- Integrating Semantic and Temporal Relationships in Facial Action Unit DetectionZhihua Li, Xiang Deng, Xiaotian Li, Lijun YinACM MM 2021 · 11 citations
- CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit RecognitionYingjie Chen, Diqi Chen, Yizhou Wang, Tao Wang et al.ACM MM 2021 · 10 citations
- Pursuing Knowledge Consistency: Supervised Hierarchical Contrastive Learning for Facial Action Unit RecognitionYingjie Chen, Chong Chen, Xiao Luo, Jianqiang Huang et al.ACM MM 2022 · 5 citations
- Facial Action Unit Intensity Estimation via Semantic Correspondence Learning with Dynamic Graph ConvolutionYingruo Fan, Jacqueline C. K. Lam, Victor On Kwok LiAAAI 2020 · 58 citations
