FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos
Yan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu, Shuyong Gao, Wei Zhang, Weifeng Ge, Wenqiang Zhang
Abstract
Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the “Happy” expression with high intensity in Talk-Show is more discriminating than the same expression with low intensity in Official-Event. To fill this gap, we build a large-scale multi-scene dataset, coined as FERV39k. We analyze the important ingredients of constructing such a novel dataset in three aspects: (1) multi-scene hierarchy and expression class, (2) generation of candidate video clips, (3) trusted manual labelling process. Based on these guidelines, we select 4 scenarios subdivided into 22 scenes, annotate 86k samples automatically obtained from 4k videos based on the well-designed workflow, and finally build 38,935 video clips labeled with 7 classic expressions. Experiment benchmarks on four kinds of baseline frame-works were also provided and further analysis on their performance across different scenes and some challenges for future research were given. Besides, we systematically investigate key components of DFER by ablation studies. The baseline framework and our project are available on https://github.com/wangyanckxx/FERV39k.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48d74290-195c-483b-bc36-8bb68f8e221bCited by top-tier papers23
- Intensity-Aware Loss for Dynamic Facial Expression Recognition in the WildHanting Li, Hongjing Niu, Zhaoqing Zhu, Feng ZhaoAAAI 2023 · 94 citations
- MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression RecognitionLicai Sun, Zheng Lian, Bin Liu, Jianhua TaoACM MM 2023 · 85 citations
- Learning Causality-inspired Representation Consistency for Video Anomaly DetectionYang Liu, Zhaoyang Xia, Mengyang Zhao, Donglai Wei et al.ACM MM 2023 · 48 citations
- Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing CorrectionAndrey V. SavchenkoICML 2023 · 43 citations
- A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party ConversationsWenjie Zheng, Jianfei Yu, Rui Xia, Shijin WangACL 2023 · 38 citations
Builds on7
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park et al.ICCV 2019 · 285 citations
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang et al.ACM MM 2020 · 205 citations
- SMART Frame Selection for Action RecognitionShreyank N. Gowda, Marcus Rohrbach, Laura Sevilla-LaraAAAI 2021 · 171 citations
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij et al.CVPR 2021
Related papers
- Rethinking the Learning Paradigm for Dynamic Facial Expression RecognitionHanyang Wang, Bo Li, Shuang Wu, Siyuan Shen et al.CVPR 2023
- OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive ProblemXinji Mai, Haoran Wang, Zeng Tao, Junxiong Lin et al.AAAI 2025
- Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual AwarenessJiaxing Zhao, Boyuan Sun, Xiang Chen, Xihan WeiAAAI 2026
- EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression RecognitionKe Wang, Yuanyuan Liu, Kejun Liu, Yuyang Xia et al.ICML 2026
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationFeng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang et al.ACM MM 2025 · 10 citations
