MIMAMO Net: Integrating Micro- and Macro-Motion for Video Emotion Recognition
Didan Deng, Zhaokang Chen, Yuqian Zhou, Bertram E. Shi
Abstract
Spatial-temporal feature learning is of vital importance for video emotion recognition. Previous deep network structures often focused on macro-motion which extends over long time scales, e.g., on the order of seconds. We believe integrating structures capturing information about both micro- and macro-motion will benefit emotion prediction, because human perceive both micro- and macro-expressions. In this paper, we propose to combine micro- and macro-motion features to improve video emotion recognition with a two-stream recurrent network, named MIMAMO (Micro-Macro-Motion) Net. Specifically, smaller and shorter micro-motions are analyzed by a two-stream network, while larger and more sustained macro-motions can be well captured by a subsequent recurrent network. Assigning specific interpretations to the roles of different parts of the network enables us to make choice of parameters based on prior knowledge: choices that turn out to be optimal. One of the important innovations in our model is the use of interframe phase differences rather than optical flow as input to the temporal stream. Compared with the optical flow, phase differences require less computation and are more robust to illumination changes. Our proposed network achieves state of the art performance on two video emotion datasets, the OMG emotion dataset and the Aff-Wild dataset. The most significant gains are for arousal prediction, for which motion information is intuitively more informative. Source code is available at https://github.com/wtomin/MIMAMO-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15c1544d-ea4d-4fb1-940f-eae11bbc5e44Cited by top-tier papers6
- Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message PassingDong Zhang, Xincheng Ju, Wei Zhang, Junhui Li et al.AAAI 2021 · 56 citations
- Contrastive Adversarial Learning for Person Independent Facial Emotion RecognitionDae Ha Kim, Byung Cheol SongAAAI 2021 · 41 citations
- Privacy-Preserving Video Classification with Convolutional Neural NetworksSikha Pentyala, Rafael Dowsley, Martine De CockICML 2021 · 25 citations
- In the Blink of an Eye: Event-based Emotion RecognitionHaiwei Zhang, Jiqing Zhang, Bo Dong, Pieter Peers et al.SIGGRAPH 2023 · 21 citations
- Optimal Transport-based Identity Matching for Identity-invariant Facial Expression RecognitionDae Ha Kim, Byung Cheol SongNeurIPS 2022 · 19 citations
Related papers
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.NeurIPS 2021 · 193 citations
- Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action RecognitionJihao Gu, Kun Li, Fei Wang, Yanyan Wei et al.ACM MM 2025 · 23 citations
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang et al.ICCV 2025 · 21 citations
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 189 citations
