EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
Dingdong WANG, Shujie LIU, Tianhua Zhang, Youjun Chen, Jinyu Li, Helen M. Meng
Abstract
Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion recognition (SER) systems, still treat emotion understanding as a simple classification problem. This provides limited interpretability of predictions, while leaving the LLMs' expressive and reasoning capabilities underutilized. In this work, we take the first step to reformulate SER as a deep reasoning problem through reinforcement learning (RL). We propose Emotion-Thinker, which is designed to generate accurate emotion predictions with interpretable explanations grounded in fine-grained acoustic cues. To achieve this, we first construct EmotionCoT-35K, an emotional reasoning dataset with Chainof-Thought annotations and detailed captions. Second, we observe that current SpeechLLMs exhibit weak prosody perception, whereas prosodic cues constitute fundamental signals for interpreting emotions. To address this, we develop the prosody-enhanced foundation model EmotionThinker-Base, and demonstrate that prosody enhancement improves emotion understanding. Third, we introduce Group-Relative-Policy-Optimization with Progressive-Trust-aware-Reasoning-Reward (GRPO-PTR) for RL. Different from standard GRPO, which relies only on rule-based outcome rewards, GRPO-PTR progressively introduces reasoning reward, dynamically adjusts it with a trustworthiness weight reflecting the alignment between reasoning and outcome, and evaluates the overall reasoning quality with a reward model based on multi-dimensional criteria. EmotionThinker outperforms previous state-of-the-art evaluation models both in emotion accuracy and explanation quality, advancing SER toward interpretable multimodal reasoning. Project page: https://github.com/dingdongwang/EmotionThinker Recent efforts have explored enhancing emotion explainability through supervised fine-tuning (SFT) (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e90cdcbb-56c7-46ab-9766-a684dc2a5d95Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankTianhe Wu, Jian Zou, Jie Liang, Lei Zhang et al.NeurIPS 2025 · 92 citations
- SECap: Speech Emotion Captioning with Large Language ModelYaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang et al.AAAI 2024 · 70 citations
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPOChengzhuo Tong, Ziyu Guo, Renrui Zhang, Wenyu Shan et al.NeurIPS 2025 · 39 citations
Related papers
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang et al.CVPR 2026 · 4 citations
- VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic AugmentationYancheng Wang, Osama Hanna, Ruiming Xie, Xianfeng Rui et al.ICLR 2026 · 4 citations
- ERCThinker: Fast-Slow Thinking for Emotion Recognition in ConversationYumeng Fu, Weitao Huang, Junjie Wu, Hao Teng et al.ACL 2026
- JoPR: Joint Emotion Perception and Reasoning for Conversational Emotion RecognitionYumeng Fu, Weitao Huang, Junjie Wu, Hao Teng et al.ACL 2026
- Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningShu Wu, Chenxing Li, Wenfu Wang, Hao Zhang et al.AAAI 2026 · 4 citations
