MIMAMO Net: Integrating Micro- and Macro-Motion for Video Emotion Recognition
Didan Deng, Zhaokang Chen, Yuqian Zhou, Bertram E. Shi
摘要
Spatial-temporal feature learning is of vital importance for video emotion recognition. Previous deep network structures often focused on macro-motion which extends over long time scales, e.g., on the order of seconds. We believe integrating structures capturing information about both micro- and macro-motion will benefit emotion prediction, because human perceive both micro- and macro-expressions. In this paper, we propose to combine micro- and macro-motion features to improve video emotion recognition with a two-stream recurrent network, named MIMAMO (Micro-Macro-Motion) Net. Specifically, smaller and shorter micro-motions are analyzed by a two-stream network, while larger and more sustained macro-motions can be well captured by a subsequent recurrent network. Assigning specific interpretations to the roles of different parts of the network enables us to make choice of parameters based on prior knowledge: choices that turn out to be optimal. One of the important innovations in our model is the use of interframe phase differences rather than optical flow as input to the temporal stream. Compared with the optical flow, phase differences require less computation and are more robust to illumination changes. Our proposed network achieves state of the art performance on two video emotion datasets, the OMG emotion dataset and the Aff-Wild dataset. The most significant gains are for arousal prediction, for which motion information is intuitively more informative. Source code is available at https://github.com/wtomin/MIMAMO-Net.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message PassingDong Zhang, Xincheng Ju, Wei Zhang, Junhui Li 等AAAI 2021 · 被引用 56 次
- Contrastive Adversarial Learning for Person Independent Facial Emotion RecognitionDae Ha Kim, Byung Cheol SongAAAI 2021 · 被引用 41 次
- Privacy-Preserving Video Classification with Convolutional Neural NetworksSikha Pentyala, Rafael Dowsley, Martine De CockICML 2021 · 被引用 25 次
- In the Blink of an Eye: Event-based Emotion RecognitionHaiwei Zhang, Jiqing Zhang, Bo Dong, Pieter Peers 等SIGGRAPH 2023 · 被引用 21 次
- Optimal Transport-based Identity Matching for Identity-invariant Facial Expression RecognitionDae Ha Kim, Byung Cheol SongNeurIPS 2022 · 被引用 19 次
相关 Paper
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 等NeurIPS 2021 · 被引用 193 次
- Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action RecognitionJihao Gu, Kun Li, Fei Wang, Yanyan Wei 等ACM MM 2025 · 被引用 23 次
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang 等ICCV 2025 · 被引用 21 次
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 被引用 189 次
