Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models
Yiyang Fang, Jian Liang, Wenke Huang, He Li, Kehua Su, Mang Ye
摘要
Multimodal large language models (MLLMs) have achieved impressive progress in tasks such as visual question answering and visual understanding, but they still face significant challenges in emotional reasoning. Current methods to enhance emotional understanding typically rely on finetuning or manual annotations, which are resourceintensive and limit scalability. In this work, we focus on improving the ability of MLLMs to capture emotions during the inference phase. Specifically, MLLMs encounter two main issues in the inference stage: they struggle to distinguish between semantically similar emotions, leading to misclassification, and they are overwhelmed by redundant or irrelevant visual information, which distracts from key emotional cues. To address these, we propose a training-free method named Sharpening Emotion Perception in MLLMs (SEPM), which incorporates a Confidence-Guided Coarseto-Fine Inference framework to refine emotion classification by guiding the model through simpler tasks. Additionally, SEPM employs Focuson-Emotion Visual Augmentation to reduce visual redundancy by directing the attention of models to relevant emotional cues in images. Experimental results demonstrate that SEPM significantly improves MLLM performance on emotionrelated tasks, providing a resource-efficient and scalable solution for emotion recognition. Our code is available in https://github.com/ fuyyyyy/SEPM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi 等NeurIPS 2025 · 被引用 39 次
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy OptimizationXuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou 等CVPR 2026 · 被引用 13 次
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang 等CVPR 2026 · 被引用 4 次
- Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment AnalysisKang He, Yuzhe Ding, Xinrong Wang, Fei Li 等CVPR 2026 · 被引用 1 次
- MASP: Multi-Aspect Guided Emotion Reasoning with Soft Prompt Tuning In Vision-Language ModelsSangEun Lee, Yubeen Lee, Eunil Park, Wonseok ChaeAAAI 2026
它引用的顶会 Paper21
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 等NeurIPS 2024 · 被引用 293 次
- EmoSet: A Large-scale Visual Emotion Dataset with Rich AttributesJingyuan Yang, Qirui Huang, Tingting Ding, Dani Lischinski 等ICCV 2023 · 被引用 111 次
- Learning from Teaching Regularization: Generalizable Correlations Should be Easy to ImitateCan Jin, Tong Che, Hongwu Peng, Yiyuan Li 等NeurIPS 2024 · 被引用 67 次
相关 Paper
- World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion RecognitionKejun Liu, Yuanyuan Liu, Ke Wang, Zhe Chen 等ICML 2026
- Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual EnhancementChongjun Tu, Peng Ye, Dongzhan Zhou, Tao Chen 等AAAI 2026
- MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMsJiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, Filip IlievskiICLR 2025 · 被引用 1 次
- Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMsXiaoyang Yi, Jing Chen, Li Peng, Yuru Bao 等ACL 2026
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma 等ICLR 2026 · 被引用 4 次
