Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the Wild
Xingjian Wang, Li Chai
摘要
In-the-wild dynamic facial expression recognition (DFER) encounters a significant challenge in recognizing emotionrelated expressions, which are often temporally and spatially diluted by emotion-irrelevant expressions and global context. Most prior DFER methods directly utilize coupled spatiotemporal representations that may incorporate weakly relevant features with emotion-irrelevant context bias. Several DFER methods highlight dynamic information for DFER, but following explicit guidance that may be vulnerable to irrelevant motion. In this paper, we propose a novel Implicit Facial Dynamics Disentanglement framework (IFDD). Through expanding wavelet lifting scheme to fully learnable framework, IFDD disentangles emotion-related dynamic information from emotion-irrelevant global context in an implicit manner, i.e., without exploit operations and external guidance. The disentanglement process contains two stages. The first is Inter-frame Static-dynamic Splitting Module (ISSM) for rough disentanglement estimation, which explores interframe correlation to generate content-aware splitting indexes on-the-fly. We utilize these indexes to split frame features into two groups, one with greater global similarity, and the other with more unique dynamic features. The second stage is Lifting-based Aggregation-Disentanglement Module (LADM) for further refinement. LADM first aggregates two groups of features from ISSM to obtain fine-grained global context features by an updater, and then disentangles emotion-related facial dynamic features from the global context by a predictor. Extensive experiments on in-the-wild datasets have demonstrated that IFDD outperforms prior supervised DFER methods with higher recognition accuracy and comparable efficiency. Code is available at https://github . com/CyberPegasus/IFDD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Robust Lightweight Facial Expression Recognition Network with Label Distribution TrainingZengqun Zhao, Qingshan Liu, Feng ZhouAAAI 2021 · 被引用 300 次
- Relative Uncertainty Learning for Facial Expression RecognitionYuhang Zhang, Chengrui Wang, Weihong DengNeurIPS 2021 · 被引用 232 次
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang 等ACM MM 2020 · 被引用 205 次
相关 Paper
- Freq-HD: An Interpretable Frequency-based High-Dynamics Affective Clip Selection Method for in-the-Wild Facial Expression Recognition in VideosZeng Tao, Yan Wang, Zhaoyu Chen, Boyang Wang 等ACM MM 2023 · 被引用 12 次
- Former-DFER: Dynamic Facial Expression Recognition TransformerZengqun Zhao, Qingshan LiuACM MM 2021 · 被引用 185 次
- Deep Disturbance-Disentangled Learning for Facial Expression RecognitionDelian Ruan, Yan Yan, Si Chen, Jing-Hao Xue 等ACM MM 2020 · 被引用 75 次
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationFeng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang 等ACM MM 2025 · 被引用 10 次
- D³Net: Dual-Branch Disturbance Disentangling Network for Facial Expression RecognitionRongyun Mo, Yan Yan, Jing-Hao Xue, Si Chen 等ACM MM 2021 · 被引用 17 次
