EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression Recognition
Ke Wang, Yuanyuan Liu, Kejun Liu, Yuyang Xia, Chang Tang, Yibing Zhan, Zhe Chen
Abstract
Dynamic Facial Expression Recognition (DFER) models the temporal evolution of facial expressions in videos. In real-world scenarios, changing scenarios distort expression trajectories, challenging existing methods. Most current approaches address this via passive feature alignment or domain-incremental learning but do not explicitly model scenario evolution, limiting their ability to capture expression dynamics under scenario-incremental changes. To address this, we propose EmWorld , an emotion world model for DFER that explicitly models latent emotion state evolution under scenario variations. Specifically, EmWorld formulates scenario-incremental DFER as a progressive Bayesian inference problem over latent world states with dual temporal scales. Slow-timescale component ( STS ) models scenario evolution using stochastic evolutionary priors, capturing long-term scenario effects and providing proactive guidance in new scenarios. Fast-timescale component ( FTS ) models frame-level expression dynamics with temporally consistent latent transitions, decoupling expression dynamics from scenario influences. By jointly inferring latent states at both timescales, EmWorld shifts DFER from a passive feature discrimination to active probabilistic state inference under evolving scenarios. Experiments on FERV39k, DFEW, and MAFW demonstrate that EmWorld consistently outperforms state-of-the-art methods, achieving up to 3.84% improvement while exhibiting strong cross-scenario stability and long-term robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e2572a9-ed61-4bef-9515-d92307d2bf19Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 397 citations
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang et al.ACM MM 2020 · 205 citations
- Former-DFER: Dynamic Facial Expression Recognition TransformerZengqun Zhao, Qingshan LiuACM MM 2021 · 185 citations
Related papers
- OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive ProblemXinji Mai, Haoran Wang, Zeng Tao, Junxiong Lin et al.AAAI 2025
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu et al.CVPR 2022 · 107 citations
- Rethinking the Learning Paradigm for Dynamic Facial Expression RecognitionHanyang Wang, Bo Li, Shuang Wu, Siyuan Shen et al.CVPR 2023
- BHGap: A Deep Iterative Prompting and Multi-stage Alignment Framework for Dynamic Facial Expression RecognitionYichi Zhang, Yunqi Han, Jiayue Ding, Liangyu ChenWWW 2026
- Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the WildXingjian Wang, Li ChaiAAAI 2025 · 1 citation
