An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos
Sicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang, Tengfei Xing, Pengfei Xu, Runbo Hu, Hua Chai, Kurt Keutzer
Abstract
Emotion recognition in user-generated videos plays an important role in human-centered computing. Existing methods mainly employ traditional two-stage shallow pipeline, i.e. extracting visual and/or audio features and training classifiers. In this paper, we propose to recognize video emotions in an end-to-end manner based on convolutional neural networks (CNNs). Specifically, we develop a deep Visual-Audio Attention Network (VAANet), a novel architecture that integrates spatial, channel-wise, and temporal attentions into a visual 3D CNN and temporal attentions into an audio 2D CNN. Further, we design a special classification loss, i.e. polarity-consistent cross-entropy loss, based on the polarity-emotion hierarchy constraint to guide the attention generation. Extensive experiments conducted on the challenging VideoEmotion-8 and Ekman-6 datasets demonstrate that the proposed VAANet outperforms the state-of-the-art approaches for video emotion recognition. Our source code is released at: https://github.com/maysonma/VAANet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 306b2ba6-664b-40d2-8591-98b0e4d0b8a8Cited by top-tier papers14
- Tailor Versatile Multi-Modal Learning for Multi-Label Emotion RecognitionYi Zhang, Mingyuan Chen, Jundong Shen, Chongjun WangAAAI 2022 · 92 citations
- Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal SpaceSicheng Zhao, Yaxian Li, Xingxu Yao, Weizhi Nie et al.ACM MM 2020 · 30 citations
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang et al.ACM MM 2023 · 29 citations
- Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker TrackingYidi Li, Hong Liu, Hao TangAAAI 2022 · 25 citations
- Privacy-Preserving Video Classification with Convolutional Neural NetworksSikha Pentyala, Rafael Dowsley, Martine De CockICML 2021 · 25 citations
Builds on1
Related papers
- Weakly Supervised Video Emotion Detection and Prediction via Cross-Modal Temporal Erasing NetworkZhicheng Zhang, Lijuan Wang, Jufeng YangCVPR 2023
- MDAN: Multi-level Dependent Attention Network for Visual Emotion AnalysisLiwen Xu, Zhengtao Wang, Bin Wu, Simon LuiCVPR 2022 · 54 citations
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 419 citations
- Cross Corpus Physiological-based Emotion Recognition Using a Learnable Visual Semantic Graph Convolutional NetworkWoan-Shiuan Chien, Hao-Chun Yang, Chi-Chun LeeACM MM 2020 · 7 citations
- STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from GaitsUttaran Bhattacharya, Trisha Mittal, Rohan Chandra, Tanmay Randhavane et al.AAAI 2020
