EmoBench: Evaluating the Emotional Intelligence of Large Language Models
Sahand Sabour, Siyang Liu, Zheyuan Zhang, June M. Liu, Jinfeng Zhou, Alvionna S. Sunaryo, Tatia M. C. Lee, Rada Mihalcea, Minlie Huang
Abstract
Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have two major shortcomings: first, they mainly focus on emotion recognition, neglecting essential EI capabilities such as emotion regulation and thought facilitation through emotion understanding; second, they are primarily constructed from existing datasets, which include frequent patterns, explicit information, and annotation errors, leading to unreliable evaluation. We propose EMOBENCH, a benchmark that draws upon established psychological theories and proposes a comprehensive definition for machine EI, including Emotional Understanding and Emotional Application. EMOBENCH includes a set of 400 hand-crafted questions in English and Chinese, which are meticulously designed to require thorough reasoning and understanding. Our findings reveal a considerable gap between the EI of existing LLMs and the average human, highlighting a promising direction for future research. Recent large language models (LLMs) (Bai et al., 043 2023; Yang et al., 2023a; Touvron et al., 2023; 044 OpenAI, 2023) have pushed the boundaries of our 045 expectations regarding their potential capabilities. 046 However, despite their apparent proficiency in a 047 variety of downstream tasks, such as question an-048 swering, and summarization (Zhou et al., 2023a; 049 Zhong et al., 2023), research on evaluating EI capa-050 bilities for LLMs has been limited. The majority of
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin et al.AAAI 2025 · 48 citations
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song et al.ACL 2025 · 12 citations
- Exploring Modular Prompt Design for Emotion and Mental Health RecognitionMinseo Kim, Taemin Kim, Thu Hoang Anh Vo, Yugyeong Jung et al.CHI 2025 · 11 citations
- Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion ReasoningMeng Luo, Bobo Li, Shanqing Xu, Shize Zhang et al.ICLR 2026 · 10 citations
- Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-PlayingWenyuan Zhang, Shuaiyi Nie, Jiawei Sheng, Zefeng Zhang et al.EMNLP 2025 · 10 citations
Builds on6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsHyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras et al.EMNLP 2023 · 21 citations
- CRoW: Benchmarking Commonsense Reasoning in Real-World TasksMete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva et al.EMNLP 2023 · 3 citations
Related papers
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language ModelsFan Zhang, Zebang Cheng, Chong Deng, Haoxuan Li et al.ICLR 2026 · 23 citations
- Consensus-Driven Multi-Agent Cognitive Reasoning for Enhancing the Emotional Intelligence of Large Language ModelsGeng Tu, Dingming Li, Jun Huang, Ruifeng XuAAAI 2026
- EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion AssessmentLancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun et al.ACM MM 2025 · 2 citations
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- MA-Bench: Towards Fine-grained Micro-Action UnderstandingKun Li, Jihao Gu, Fei Wang, Zhiliang Wu et al.CVPR 2026 · 12 citations
