Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
Daiqing Wu, Dongbao Yang, Yu Zhou, Can Ma
Abstract
Visual emotion recognition (VER) is a longstanding field that has garnered increasing attention with the advancement of deep neural networks. Although recent studies have achieved notable improvements by leveraging the knowledge embedded within pre-trained visual models, the lack of direct association between factual-level features and emotional categories, called the ''affective gap'', limits the applicability of pre-training knowledge for VER tasks. On the contrary, the explicit emotional expression and high information density in textual modality eliminate the ''affective gap''. Therefore, we propose borrowing the knowledge from the pre-trained textual model to enhance the emotional perception of pre-trained visual models. We focus on the factual and emotional connections between images and texts in noisy social media data, and propose Partitioned Adaptive Contrastive Learning (PACL) to leverage these connections. Specifically, we manage to separate different types of samples and devise distinct contrastive learning strategies for each type. By dynamically constructing negative and positive pairs, we fully exploit the potential of noisy samples. Through comprehensive experiments, we demonstrate that bridging the "affective gap'' significantly improves the performance of various pre-trained visual models in downstream emotion-related tasks. Our code is released on https://github.com/wdqqdw/PACL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eaf1380a-3fb4-4c50-aed0-c165374f038dCited by top-tier papers3
- Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go BeyondHuiyu Zhai, Xingxing Yang, Yalan Ye, Chenyang Li et al.ACM MM 2025 · 5 citations
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICLR 2026 · 4 citations
- An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception CapabilityDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICML 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
Related papers
- Progressive Visual Content Understanding Network for Image Emotion ClassificationJicai Pan, Shangfei WangACM MM 2023 · 5 citations
- Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive LearningJishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang et al.CVPR 2023
- Semantic-Augmented Image Clustering via Adaptive Multi-Modal CollaborationXiaohan Zhang, Chao Zhang, Deng Xu, Hong Yu et al.AAAI 2026
- CMAL: A Novel Cross-Modal Associative Learning Framework for Vision-Language Pre-TrainingZhiyuan Ma, Jianjun Li, Guohui Li, Kaiyan HuangACM MM 2022 · 7 citations
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu et al.ACM MM 2025 · 9 citations
