Multiple Feature Refining Network for Visual Emotion Distribution Learning
Qinfu Xu, Shaozu Yuan, Yiwei Wei, Jie Wu, Leiquan Wang, Chunlei Wu
Abstract
The significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle emotion awareness. To learn the distribution of emotions involved in images, most previous works learn coarse semantic knowledge with unbiased filtering. Consequently, they focus on the entire scene and suffer from the redundancy of semantic-irrelevant information, which diminishes the affective coherence, impeding the comprehension of emotional attributes within the treated features. In light of this, we reanalyze from the perspective of information filtering and propose a novel method called Multiple Feature Refining Network (MFRN). To minimize low-level feature redundancy, we design a wavelet-based separated frequency modeling, named Spectral Mixer, to learn invariant representations and enhance emotion saliency in low-level image features. At the higher semantic level, we design a Semantic Graph Prompt Learning for emotional semantic filtering, ensuring the purity of emotional information and providing the model with richer content semantics. Experiments conducted on three commonly used datasets have demonstrated the superiority of our MFRN model over cutting-edge methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03060ffa-e77d-4368-8991-190ee96b5f5cBuilds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- Bridging the Gap Between Vision Transformers and Convolutional Neural Networks on Small DatasetsZhiying Lu, Hongtao Xie, Chuanbin Liu, Yongdong ZhangNeurIPS 2022 · 107 citations
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang et al.ICLR 2023 · 62 citations
- Scattering Vision Transformer: Spectral Mixing MattersBadri N. Patro, Vijay AgneeswaranNeurIPS 2023 · 43 citations
Related papers
- StyleEDL: Style-Guided High-order Attention Network for Image Emotion Distribution LearningPeiguang Jing, Xianyi Liu, Ji Wang, Yinwei Wei et al.ACM MM 2023 · 5 citations
- A Circular-Structured Representation for Visual Emotion Distribution LearningJingyuan Yang, Jie Li, Leida Li, Xiumei Wang et al.CVPR 2021
- MDAN: Multi-level Dependent Attention Network for Visual Emotion AnalysisLiwen Xu, Zhengtao Wang, Bin Wu, Simon LuiCVPR 2022 · 54 citations
- HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution LearningChuhang Zheng, Chunwei Tian, Jie Wen, Daoqiang Zhang et al.ACM MM 2025 · 13 citations
- Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video CaptioningCheng Ye, Weidong Chen, Peipei Song, Xinyan Liu et al.ACM MM 2025 · 9 citations
