From Pixels to Semantics: Unified Facial Action Representation Learning for Micro-Expression Analysis
Yicheng Deng, Hideaki Hayashi, Hajime Nagahara
Abstract
Micro-expression recognition (MER) is highly challenging due to the subtle and rapid facial muscle movements and the scarcity of annotated data. Existing methods typically rely on pixel-level motion descriptors such as optical flow and frame difference, which tend to be sensitive to identity and lack generalization. In this work, we propose D-FACE, a Discrete Facial ACtion Encoding framework that leverages large-scale facial video data to pretrain an identity- and domain-invariant facial action tokenizer, for MER. For the first time, MER is shifted from relying on pixel-level motion descriptors to leveraging semantic-level facial action tokens, providing compact and generalizable representations of facial dynamics. Empirical analyses reveal that these tokens exhibit position-dependent semantics, motivating sequential modeling. Building on this insight, we employ a Transformer with sparse attention pooling to selectively capture discriminative action cues. Furthermore, to explicitly bridge action tokens with human-understandable emotions, we introduce an emotion-description-guided CLIP (EDCLIP) alignment. EDCLIP leverages textual prompts as semantic anchors for representation learning, while enforcing that the "others" category, which lacks corresponding prompts due to its ambiguity, remains distant from all anchor prompts. Extensive experiments on multiple datasets demonstrate that our method achieves not only state-of-the-art recognition accuracy but also high-quality cross-identity and even cross-domain micro-expression generation, suggesting a paradigm shift from pixel-level to generalizable semantic-level facial motion analysis. Code is available at https://github.com/KinopioIsAllIn/D-FACE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb73fb58-f01a-44be-b6b5-623a43dcb60cBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- A Novel Graph-TCN with a Graph Structured Representation for Micro-expression RecognitionLing Lei, Jianfeng Li, Tong Chen, Shigang LiACM MM 2020 · 134 citations
- FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERsHaodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng et al.ACM MM 2024 · 26 citations
- Latent Action Pretraining from VideosSeonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo et al.ICLR 2025
Related papers
- Domain Knowledge Enhanced Vision-Language Pretrained Model for Dynamic Facial Expression RecognitionLiupeng Li, Yuhua Zheng, Shupeng Liu, Xiaoyin Xu et al.ACM MM 2024 · 4 citations
- Multimodal Prompt Alignment for Facial Expression RecognitionFuyan Ma, Yiran He, Bin Sun, Shutao LiICCV 2025 · 5 citations
- SelfME: Self-Supervised Motion Learning for Micro-Expression RecognitionXinqi Fan, Xueli Chen, Mingjie Jiang, Ali Raza Shahid et al.CVPR 2023
- CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit RecognitionYingjie Chen, Diqi Chen, Yizhou Wang, Tao Wang et al.ACM MM 2021 · 10 citations
- Feature Representation Learning with Adaptive Displacement Generation and Transformer Fusion for Micro-Expression RecognitionZhijun Zhai, Jianhui Zhao, Chengjiang Long, Wenju Xu et al.CVPR 2023
