Progressive Visual Content Understanding Network for Image Emotion Classification
Jicai Pan, Shangfei Wang
Abstract
Most existing methods for image emotion classification extract features directly from images supervised by a single emotional label. However, this approach has a limitation known as the affective gap which restricts the capability of these features as they do not always align with the emotions perceived by users. To effectively bridge the affective gap, this paper proposes a visual content understanding network inspired by the human staged emotion perception process. The proposed network is comprised of three perception modules designed to extract multi-level information. Firstly, an entity perception module extracts entities from images. Secondly, an attribute perception module extracts the attribute content of each entity. Thirdly, an emotion perception module extracts emotion features based on both the entity and attribute information. We generate pseudo-labels of entities and attributes through image segmentation and vision-language models to provide auxiliary guidance for network learning. The progressive entity and attribute understanding enable the network to hierarchically extract semantic-level features for emotion analysis. Extensive experiments demonstrate that our progressive learning network achieves superior performance on various benchmark datasets for image emotion classification.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ccc66bce-ae89-4117-80d8-1f5cb573be6cCited by top-tier papers2
- To Err Like Human: Affective Bias-Inspired Measures for Visual Emotion Recognition EvaluationChenxi Zhao, Jinglei Shi, Liqiang Nie, Jufeng YangNeurIPS 2024 · 10 citations
- Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go BeyondHuiyu Zhai, Xingxing Yang, Yalan Ye, Chenyang Li et al.ACM MM 2025 · 5 citations
Related papers
- MDAN: Multi-level Dependent Attention Network for Visual Emotion AnalysisLiwen Xu, Zhengtao Wang, Bin Wu, Simon LuiCVPR 2022 · 54 citations
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang et al.ACM MM 2023 · 29 citations
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text PairsDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 6 citations
- E-CORE: Emotion Correlation Enhanced Empathetic Dialogue GenerationFengyi Fu, Lei Zhang, Quan Wang, Zhendong MaoEMNLP 2023 · 8 citations
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu et al.ACM MM 2025 · 9 citations
