Multiple Feature Refining Network for Visual Emotion Distribution Learning
Qinfu Xu, Shaozu Yuan, Yiwei Wei, Jie Wu, Leiquan Wang, Chunlei Wu
摘要
The significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle emotion awareness. To learn the distribution of emotions involved in images, most previous works learn coarse semantic knowledge with unbiased filtering. Consequently, they focus on the entire scene and suffer from the redundancy of semantic-irrelevant information, which diminishes the affective coherence, impeding the comprehension of emotional attributes within the treated features. In light of this, we reanalyze from the perspective of information filtering and propose a novel method called Multiple Feature Refining Network (MFRN). To minimize low-level feature redundancy, we design a wavelet-based separated frequency modeling, named Spectral Mixer, to learn invariant representations and enhance emotion saliency in low-level image features. At the higher semantic level, we design a Semantic Graph Prompt Learning for emotional semantic filtering, ensuring the purity of emotional information and providing the model with richer content semantics. Experiments conducted on three commonly used datasets have demonstrated the superiority of our MFRN model over cutting-edge methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- Bridging the Gap Between Vision Transformers and Convolutional Neural Networks on Small DatasetsZhiying Lu, Hongtao Xie, Chuanbin Liu, Yongdong ZhangNeurIPS 2022 · 被引用 107 次
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang 等ICLR 2023 · 被引用 62 次
- Scattering Vision Transformer: Spectral Mixing MattersBadri N. Patro, Vijay AgneeswaranNeurIPS 2023 · 被引用 43 次
相关 Paper
- StyleEDL: Style-Guided High-order Attention Network for Image Emotion Distribution LearningPeiguang Jing, Xianyi Liu, Ji Wang, Yinwei Wei 等ACM MM 2023 · 被引用 5 次
- A Circular-Structured Representation for Visual Emotion Distribution LearningJingyuan Yang, Jie Li, Leida Li, Xiumei Wang 等CVPR 2021
- MDAN: Multi-level Dependent Attention Network for Visual Emotion AnalysisLiwen Xu, Zhengtao Wang, Bin Wu, Simon LuiCVPR 2022 · 被引用 54 次
- HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution LearningChuhang Zheng, Chunwei Tian, Jie Wen, Daoqiang Zhang 等ACM MM 2025 · 被引用 13 次
- Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video CaptioningCheng Ye, Weidong Chen, Peipei Song, Xinyan Liu 等ACM MM 2025 · 被引用 9 次
