MDAN: Multi-level Dependent Attention Network for Visual Emotion Analysis
Liwen Xu, Zhengtao Wang, Bin Wu, Simon Lui
Abstract
Visual Emotion Analysis (VEA) is attracting increasing attention. One of the biggest challenges of VEA is to bridge the affective gap between visual clues in a picture and the emotion expressed by the picture. As the granularity of emotions increases, the affective gap increases as well. Existing deep approaches try to bridge the gap by directly learning discrimination among emotions globally in one shot. They ignore the hierarchical relationship among emotions at different affective levels, and the variation in the affective level of emotions to be classified. In this paper, we present the multi-level dependent attention network (MDAN) with two branches to leverage the emotion hierarchy and the correlation between different affective levels and semantic levels. The bottom-up branch directly learns emotions at the highest affective level and largely prevents hierarchy violation by explicitly following the emotion hierarchy while predicting emotions at lower affective levels. In contrast, the top-down branch aims to disentangle the affective gap by one-to-one mapping between semantic levels and affective levels, namely, Affective Semantic Mapping. A local classifier is appended at each semantic level to learn discrimination among emotions at the corresponding affective level. Then, we integrate global learning and local learning into a unified deep framework and optimize it simultaneously. Moreover, to properly model channel dependencies and spatial attention while disentangling the affective gap, we carefully designed two attention modules: the Multi-head Cross Channel Attention module and the Level-dependent Class Activation Map module. Finally, the proposed deep framework obtains new state-of-the-art performance on six VEA benchmarks, where it outperforms existing state-of-the-art methods by a large margin, e.g., +3.85% on the WEBEmo dataset at 25 classes classification accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce0c7d38-f40f-4eae-ac8a-a109a52c390aCited by top-tier papers13
- EmoSet: A Large-scale Visual Emotion Dataset with Rich AttributesJingyuan Yang, Qirui Huang, Tingting Ding, Dani Lischinski et al.ICCV 2023 · 111 citations
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang et al.ACM MM 2023 · 29 citations
- EmoVIT: Revolutionizing Emotion Insights with Visual Instruction TuningHongxia Xie, Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen et al.CVPR 2024 · 21 citations
- Dual-path Collaborative Generation Network for Emotional Video CaptioningCheng Ye, Weidong Chen, Jingyu Li, Lei Zhang et al.ACM MM 2024 · 15 citations
- To Err Like Human: Affective Bias-Inspired Measures for Visual Emotion Recognition EvaluationChenxi Zhao, Jinglei Shi, Liqiang Nie, Jufeng YangNeurIPS 2024 · 10 citations
Builds on2
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bottleneck Transformers for Visual RecognitionAravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens et al.CVPR 2021
Related papers
- Progressive Visual Content Understanding Network for Image Emotion ClassificationJicai Pan, Shangfei WangACM MM 2023 · 5 citations
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu et al.ACM MM 2025 · 9 citations
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosSicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang et al.AAAI 2020 · 123 citations
- Deepfake Video Detection via Facial Action Dependencies EstimationLingfeng Tan, Yunhong Wang, Junfu Wang, Liang Yang et al.AAAI 2023 · 29 citations
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 11 citations
