ReactioNet: Learning High-order Facial Behavior from Universal Stimulus-Reaction by Dyadic Relation Reasoning
Xiaotian Li, Taoyue Wang, Geran Zhao, Xiang Zhang, Xi Kang, Lijun Yin
Abstract
Diverse visual stimuli can evoke various human affective states, which are usually manifested in an individual’s muscular actions and facial expressions. In lab-controlled emotion datasets, such a critical component (i.e., stimulus) was commonly designed in a limited way, making researchers incapable of generalizing the universal correlation and causation of stimulus-reaction as well as predicting possible emotions from context, timing, and relation. In this paper, we collected a large-scale spontaneous facial behavior database ReactioNet, which contains 1.1 million coupled stimulus-reaction tuples (visual/audio/caption from both stimuli and subjects). We introduce a new facial behavior detection scenario, Dyadic Relation Reasoning (DRR), which aims to detect facial actions by reasoning their relations with stimuli. By aggregating the dyadic information, our method essentially forms a relation prototype Universal Stimulus Reaction (U-SR), which encodes the low-order and high-order relationships between stimulus agents and facial reactions. A framework with both non-graph and graph modules is further developed to evaluate DRR-based facial action unit detection, facial expression recognition, and scene classification. Specifically, to learn "what" arouses a facial reaction, the non-graph module associates and projects the fine-grained stimulus-reaction features into common subspaces using cross-domain contrastive learning. To learn "how" stimulus-reaction pairs are mutually related, the graph module adopts Graph Convolution Network to represent, converge, and infer the dyadic U-SR relation under two relation assumptions (i.e., homophily and heterophily [68]). Extensive experiments demonstrate the effectiveness of the proposed work. The new dataset will be available for the research community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Multi-Label Compound Expression Recognition: C-EXPR Database & NetworkDimitrios KolliasCVPR 2023
- ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion ModelCheng Luo, Siyang Song, Siyuan Yan, Zhen Yu et al.ACM MM 2025 · 1 citation
- Knowledge Augmented Deep Neural Networks for Joint Facial Expression and Action Unit RecognitionZijun Cui, Tengfei Song, Yuru Wang, Qiang JiNeurIPS 2020 · 70 citations
- Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior UnderstandingXiang Zhang, Taoyue Wang, Xiaotian Li, Huiyuan Yang et al.ICCV 2023 · 26 citations
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang et al.ACM MM 2020 · 205 citations
