EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
Li Zhou, Lutong Yu, You Lyu, Yihang Lin, Zefeng Zhao, Junyi Ao, Yuhao Zhang, Benyou Wang, Haizhou Li
摘要
Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks typically evaluate linguistic, acoustic, reasoning, or dialogue abilities in isolation, overlooking the integration of these skills that is crucial for human-like, emotionally intelligent conversation. We present EchoMind, the first interrelated, multi-level benchmark that simulates the cognitive process of empathetic dialogue through sequential, context-linked tasks: spoken-content understanding, vocal-cue perception, integrated reasoning, and response generation. All tasks share identical and semantically neutral scripts that are free of explicit emotional or contextual cues, and controlled variations in vocal style are used to test the effect of delivery independent of the transcript. EchoMind is grounded in an empathy-oriented framework spanning 3 coarse and 12 fine-grained dimensions, encompassing 39 vocal attributes, and evaluated using both objective and subjective metrics. Testing 12 advanced SLMs reveals that even state-of-the-art models struggle with highexpressive vocal cues, limiting empathetic response quality. Analyses of prompt strength, speech source, and ideal vocal cue recognition reveal persistent weaknesses in instruction-following, resilience to natural speech variability, and effective use of vocal cues for empathy. These results underscore the need for SLMs that integrate linguistic content with diverse vocal cues to achieve truly empathetic conversational ability. Project website: https://hlt-cuhksz.github.io/EchoMind/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar 等NeurIPS 2025 · 被引用 299 次
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang 等ICLR 2026 · 被引用 143 次
- Audio Entailment: Assessing Deductive Reasoning for Audio UnderstandingSoham Deshmukh, Shuo Han, Hazim T. Bukhari, Benjamin Elizalde 等AAAI 2025 · 被引用 23 次
- Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken ConversationsGuan-Ting Lin, Cheng-Han Chiang, Hung-yi LeeACL 2024 · 被引用 15 次
相关 Paper
- EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion AssessmentLancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun 等ACM MM 2025 · 被引用 2 次
- Consensus-Driven Multi-Agent Cognitive Reasoning for Enhancing the Emotional Intelligence of Large Language ModelsGeng Tu, Dingming Li, Jun Huang, Ruifeng XuAAAI 2026
- HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech UnderstandingChen Li, Peiji Yang, Yicheng Zhong, Jianxing Yu 等AAAI 2026 · 被引用 1 次
- MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal InteractionsRamaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi 等EMNLP 2025
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language ModelsFan Zhang, Zebang Cheng, Chong Deng, Haoxuan Li 等ICLR 2026 · 被引用 23 次
