DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in Videos
Yan Wang, Yixuan Sun, Wei Song, Shuyong Gao, Yiwen Huang, Zhaoyu Chen, Weifeng Ge, Wenqiang Zhang
Abstract
Current works of facial expression learning in video consume significant computational resources to learn spatial channel feature representations and temporal relationships. To mitigate this issue, we propose a Dual Path multi-excitation Collaborative Network (DPCNet) to learn the critical information for facial expression representation from fewer keyframes in videos. Specifically, the DPCNet learns the important regions and keyframes from a tuple of four view-grouped frames by multi-excitation modules and produces dual-path representations of one video with consistency under two regularization strategies. A spatial-frame excitation module and a channel-temporal aggregation module are introduced consecutively to learn spatial-frame representation and generate complementary channel-temporal aggregation, respectively. Moreover, we design a multi-frame regularization loss to enforce the representation of multiple frames in the dual view to be semantically coherent. To obtain consistent prediction probabilities from the dual path, we further propose a dual path regularization loss, aiming to minimize the divergence between the distributions of two-path embeddings. Extensive experiments and ablation studies show that the DPCNet can significantly improve the performance of video-based FER and achieve state-of-the-art results on the large-scale DFEW dataset.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 80ac06b4-b8f1-4349-a4f1-b4a31aebffb7Cited by top-tier papers9
- MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression RecognitionLicai Sun, Zheng Lian, Bin Liu, Jianhua TaoACM MM 2023 · 85 citations
- Learning Causality-inspired Representation Consistency for Video Anomaly DetectionYang Liu, Zhaoyang Xia, Mengyang Zhao, Donglai Wei et al.ACM MM 2023 · 48 citations
- OpenVIS: Open-vocabulary Video Instance SegmentationPinxue Guo, Hao Huang, Peiyang He, Xuefeng Liu et al.AAAI 2025 · 26 citations
- FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERsHaodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng et al.ACM MM 2024 · 26 citations
- All rivers run into the sea: Unified Modality Brain-Inspired Emotional Central MechanismXinji Mai, Junxiong Lin, Haoran Wang, Zeng Tao et al.ACM MM 2024 · 7 citations
Related papers
- Rethinking the Learning Paradigm for Dynamic Facial Expression RecognitionHanyang Wang, Bo Li, Shuang Wu, Siyuan Shen et al.CVPR 2023
- Patch-Aware Representation Learning for Facial Expression RecognitionYi Wu, Shangfei Wang, Yanan ChangACM MM 2023 · 5 citations
- Former-DFER: Dynamic Facial Expression Recognition TransformerZengqun Zhao, Qingshan LiuACM MM 2021 · 185 citations
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang et al.ACM MM 2020 · 205 citations
- A Brain-Inspired Way of Reducing the Network Complexity via Concept-Regularized Coding for Emotion RecognitionHan Lu, Xiahai Zhuang, Qiang LuoAAAI 2024
