EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression Recognition
Ke Wang, Yuanyuan Liu, Kejun Liu, Yuyang Xia, Chang Tang, Yibing Zhan, Zhe Chen
摘要
Dynamic Facial Expression Recognition (DFER) models the temporal evolution of facial expressions in videos. In real-world scenarios, changing scenarios distort expression trajectories, challenging existing methods. Most current approaches address this via passive feature alignment or domain-incremental learning but do not explicitly model scenario evolution, limiting their ability to capture expression dynamics under scenario-incremental changes. To address this, we propose EmWorld , an emotion world model for DFER that explicitly models latent emotion state evolution under scenario variations. Specifically, EmWorld formulates scenario-incremental DFER as a progressive Bayesian inference problem over latent world states with dual temporal scales. Slow-timescale component ( STS ) models scenario evolution using stochastic evolutionary priors, capturing long-term scenario effects and providing proactive guidance in new scenarios. Fast-timescale component ( FTS ) models frame-level expression dynamics with temporally consistent latent transitions, decoupling expression dynamics from scenario influences. By jointly inferring latent states at both timescales, EmWorld shifts DFER from a passive feature discrimination to active probabilistic state inference under evolving scenarios. Experiments on FERV39k, DFEW, and MAFW demonstrate that EmWorld consistently outperforms state-of-the-art methods, achieving up to 3.84% improvement while exhibiting strong cross-scenario stability and long-term robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 被引用 397 次
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang 等ACM MM 2020 · 被引用 205 次
- Former-DFER: Dynamic Facial Expression Recognition TransformerZengqun Zhao, Qingshan LiuACM MM 2021 · 被引用 185 次
相关 Paper
- OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive ProblemXinji Mai, Haoran Wang, Zeng Tao, Junxiong Lin 等AAAI 2025
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu 等CVPR 2022 · 被引用 107 次
- Rethinking the Learning Paradigm for Dynamic Facial Expression RecognitionHanyang Wang, Bo Li, Shuang Wu, Siyuan Shen 等CVPR 2023
- BHGap: A Deep Iterative Prompting and Multi-stage Alignment Framework for Dynamic Facial Expression RecognitionYichi Zhang, Yunqi Han, Jiayue Ding, Liangyu ChenWWW 2026
- Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the WildXingjian Wang, Li ChaiAAAI 2025 · 被引用 1 次
