VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
Yancheng Wang, Osama Hanna, Ruiming Xie, Xianfeng Rui, Maohao Shen, Xuedong Zhang, Christian Fuegen, Jilong Wu, Debjyoti Paul, Arthur Guo, Zhihong Lei, Ozlem Kalinli
Abstract
Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal dynamics. Although large language models (LLMs) have shown promise in reasoning over textual transcriptions for emotion recognition, they typically neglect fine-grained prosodic information, limiting their effectiveness and interpretability. In this work, we propose VowelPrompt, a linguistically grounded framework that augments LLM-based emotion recognition with interpretable, fine-grained vowel-level prosodic cues. Drawing on phonetic evidence that vowels serve as primary carriers of affective prosody, VowelPrompt extracts pitch-, energy-, and duration-based descriptors from time-aligned vowel segments, and converts these features into natural language descriptions for better interpretability. Such a design enables LLMs to jointly reason over lexical semantics and fine-grained prosodic variation. Moreover, we adopt a two-stage adaptation procedure comprising supervised fine-tuning (SFT) followed by Reinforcement Learning with Verifiable Reward (RLVR), implemented via Group Relative Policy Optimization (GRPO), to enhance reasoning capability, enforce structured output adherence, and improve generalization across domains and speaker variations. Extensive evaluations across diverse benchmark datasets demonstrate that VowelPrompt consistently outperforms state-of-the-art emotion recognition methods under zero-shot, fine-tuned, cross-domain, and cross-linguistic conditions, while enabling the generation of interpretable explanations that are jointly grounded in contextual semantics and fine-grained prosodic structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang et al.NeurIPS 2024 · 293 citations
Related papers
- EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion ReasoningDingdong WANG, Shujie LIU, Tianhua Zhang, Youjun Chen et al.ICLR 2026 · 22 citations
- ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools — From Consensus Learning to Ambiguity-Driven Emotion ReasoningEsther Sun, Bo-Hao Su, Abinay Reddy Naini, Shinji Watanabe et al.ICML 2026 · 1 citation
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang et al.CVPR 2026 · 4 citations
- Visual Prompting in LLMs for Enhancing Emotion RecognitionQixuan Zhang, Zhifeng Wang, Dylan Zhang, Wenjia Niu et al.EMNLP 2024 · 5 citations
- MASP: Multi-Aspect Guided Emotion Reasoning with Soft Prompt Tuning In Vision-Language ModelsSangEun Lee, Yubeen Lee, Eunil Park, Wonseok ChaeAAAI 2026
