STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
Uttaran Bhattacharya, Trisha Mittal, Rohan Chandra, Tanmay Randhavane, Aniket Bera, Dinesh Manocha
Abstract
We present a novel classifier network called STEP, to classify perceived human emotion from gaits, based on a Spatial Temporal Graph Convolutional Network (ST-GCN) architecture. Given an RGB video of an individual walking, our formulation implicitly exploits the gait features to classify the emotional state of the human into one of four emotions: happy, sad, angry, or neutral. We use hundreds of annotated real-world gait videos and augment them with thousands of annotated synthetic gaits generated using a novel generative network called STEP-Gen, built on an ST-GCN based Conditional Variational Autoencoder (CVAE). We incorporate a novel push-pull regularization loss in the CVAE formulation of STEP-Gen to generate realistic gaits and improve the classification accuracy of STEP. We also release a novel dataset (E-Gait), which consists of 2, 177 human gaits annotated with perceived emotions along with thousands of synthetic gaits. In practice, STEP can learn the affective features and exhibits classification accuracy of 89% on E-Gait, which is 14-30% more accurate over prior methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d0399f1-2e3a-4d3f-b3a6-6b0344e1df52Cited by top-tier papers16
- Emotions Don't Lie: An Audio-Visual Deepfake Detection Method using Affective CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera et al.ACM MM 2020 · 314 citations
- Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2gUttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan et al.IEEE VR 2021 · 147 citations
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression LearningUttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh ManochaACM MM 2021 · 94 citations
- Leveraging Activity Recognition to Enable Protective Behavior Detection in Continuous DataChongyang Wang, Yuan Gao, Akhil Mathur, Amanda C. de C. Williams et al.UbiComp 2021 · 43 citations
- Temporal Segmentation of Fine-gained Semantic Action: A Motion-Centered Figure Skating DatasetShenglan Liu, Aibin Zhang, Yunheng Li, Jian Zhou et al.AAAI 2021 · 33 citations
Related papers
- Multimodal Adaptive Emotion Transformer with Flexible Modality Inputs on A Novel Dataset with Continuous LabelsWei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang LuACM MM 2023 · 44 citations
- GANmut: Learning Interpretable Conditional Space for Gamut of EmotionsStefano d'Apolito, Danda Pani Paudel, Zhiwu Huang, Andrés Romero et al.CVPR 2021
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosSicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang et al.AAAI 2020 · 123 citations
- EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's PrincipleTrisha Mittal, Pooja Guhan, Uttaran Bhattacharya, Rohan Chandra et al.CVPR 2020
- Message Passing on Semantic-Anchor-Graphs for Fine-grained Emotion Representation Learning and ClassificationPinyi Zhang, Jingyang Chen, Junchen Shen, Zijie Zhai et al.EMNLP 2024 · 1 citation
