Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models
Yiyang Fang, Jian Liang, Wenke Huang, He Li, Kehua Su, Mang Ye
Abstract
Multimodal large language models (MLLMs) have achieved impressive progress in tasks such as visual question answering and visual understanding, but they still face significant challenges in emotional reasoning. Current methods to enhance emotional understanding typically rely on finetuning or manual annotations, which are resourceintensive and limit scalability. In this work, we focus on improving the ability of MLLMs to capture emotions during the inference phase. Specifically, MLLMs encounter two main issues in the inference stage: they struggle to distinguish between semantically similar emotions, leading to misclassification, and they are overwhelmed by redundant or irrelevant visual information, which distracts from key emotional cues. To address these, we propose a training-free method named Sharpening Emotion Perception in MLLMs (SEPM), which incorporates a Confidence-Guided Coarseto-Fine Inference framework to refine emotion classification by guiding the model through simpler tasks. Additionally, SEPM employs Focuson-Emotion Visual Augmentation to reduce visual redundancy by directing the attention of models to relevant emotional cues in images. Experimental results demonstrate that SEPM significantly improves MLLM performance on emotionrelated tasks, providing a resource-efficient and scalable solution for emotion recognition. Our code is available in https://github.com/ fuyyyyy/SEPM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi et al.NeurIPS 2025 · 39 citations
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy OptimizationXuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou et al.CVPR 2026 · 13 citations
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang et al.CVPR 2026 · 4 citations
- Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment AnalysisKang He, Yuzhe Ding, Xinrong Wang, Fei Li et al.CVPR 2026 · 1 citation
- MASP: Multi-Aspect Guided Emotion Reasoning with Soft Prompt Tuning In Vision-Language ModelsSangEun Lee, Yubeen Lee, Eunil Park, Wonseok ChaeAAAI 2026
Builds on21
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang et al.NeurIPS 2024 · 293 citations
- EmoSet: A Large-scale Visual Emotion Dataset with Rich AttributesJingyuan Yang, Qirui Huang, Tingting Ding, Dani Lischinski et al.ICCV 2023 · 111 citations
- Learning from Teaching Regularization: Generalizable Correlations Should be Easy to ImitateCan Jin, Tong Che, Hongwu Peng, Yiyuan Li et al.NeurIPS 2024 · 67 citations
Related papers
- World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion RecognitionKejun Liu, Yuanyuan Liu, Ke Wang, Zhe Chen et al.ICML 2026
- Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual EnhancementChongjun Tu, Peng Ye, Dongzhan Zhou, Tao Chen et al.AAAI 2026
- MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMsJiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, Filip IlievskiICLR 2025 · 1 citation
- Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMsXiaoyang Yi, Jing Chen, Li Peng, Yuru Bao et al.ACL 2026
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICLR 2026 · 4 citations
