Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
Hyeongseop Rha, Jeong Hun Yeo, Yeonju Kim, Yong Man Ro
摘要
The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize this shift, emotion understanding becomes essential allowing systems to capture subtle cues underlying user intent. Furthermore, providing faithful explanations for predicted emotions is crucial to ensure interpretability and build user trust. However, current MLLM-based methods often generate emotion explanations that diverge from the target labels and sometimes even contradict their own predicted emotions. This inconsistency poses a critical risk for misunderstanding and erodes reliability in interactive settings. To address this, we propose a novel approach: the Emotional Rationale Verifier (ERV) and an Explanation Reward. Our method guides the model to produce reasoning that is explicitly consistent with the target emotion during multimodal emotion recognition without modifying the model architecture or requiring additional paired video-description annotations. Our method significantly improves faithful explanation-prediction consistency and explanation emotion accuracy on the MAFW and DFEW datasets. Through extensive experiments and human evaluations, we show that our approach not only enhances alignment between explanation and prediction but also empowers MLLMs to deliver emotionally coherent, trustworthy interactions, marking a key step toward truly human-like HCI systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park 等ICCV 2019 · 被引用 285 次
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang 等ACM MM 2020 · 被引用 205 次
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu 等CVPR 2022 · 被引用 107 次
- MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression RecognitionLicai Sun, Zheng Lian, Bin Liu, Jianhua TaoACM MM 2023 · 被引用 85 次
相关 Paper
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang 等CVPR 2026 · 被引用 4 次
- DEEMO: De-identity Multimodal Emotion Recognition and ReasoningDeng Li, Bohao Xing, Xin Liu, Baiqiang Xia 等ACM MM 2025 · 被引用 8 次
- ContextFace: Generating Facial Expressions from Emotional ContextsMin-Jung Kim, Minsang Kim, Seung Jun BaekICCV 2025 · 被引用 1 次
- Causal-ERC: A Multimodal Framework with Causal Prompting for Emotion Recognition in Conversations with Large Language ModelsRan Jing, Geng Tu, Yice Zhang, Ruifeng XuAAAI 2026
- Consensus-Driven Multi-Agent Cognitive Reasoning for Enhancing the Emotional Intelligence of Large Language ModelsGeng Tu, Dingming Li, Jun Huang, Ruifeng XuAAAI 2026
