Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language Models
Nicholas Deas, Iván Ernesto Pérez Mejía, Ellie Yang, Kathleen McKeown
Abstract
Prior work evaluating emotion and affective understanding in large language models (LLMs) typically rely on predetermined label sets or focus on a singular evaluation task (e.g., emotion detection). We consider affective states, referring to the much broader variety of terms people use to label their emotional experiences. We evaluate multilingual language models' understanding of affective states in English and Spanish through three different tasks: 1) identification, where models predict an affective state given text, 2) expression, where models generate text expressing a given affective state, and 3) verification, where models report whether a given term refers to an affective state. We show that performance on one task is not necessarily predictive of performance on another. Using these three tasks, we then begin to explore when and why models struggle to understand particular affective states compared to others. We examine systematic patterns in the affective state terms that are well and poorly understood by models, characterizing the working emotion vocabulary of LLMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8845036-e8b7-45f6-a386-b91f525c4bf7Builds on8
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia et al.ACL 2024 · 52 citations
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen et al.ACL 2020 · 16 citations
- MASIVE: Open-Ended Affective State Identification in English and SpanishNicholas Deas, Elsbeth Turcan, Iván Pérez Mejía, Kathleen R. McKeownEMNLP 2024 · 3 citations
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-IndexHao Xu, Jiacheng Liu, Yejin Choi, Noah A. Smith et al.EMNLP 2025 · 1 citation
Related papers
- CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion UnderstandingTadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom II, Iqra Ameer et al.ACL 2025
- Exploring Modular Prompt Design for Emotion and Mental Health RecognitionMinseo Kim, Taemin Kim, Thu Hoang Anh Vo, Yugyeong Jung et al.CHI 2025 · 11 citations
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICLR 2026 · 4 citations
- F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsYu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li et al.ACL 2024
- Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language ModelsBenjamin Z. Reichman, Adar Avsian, Larry HeckICLR 2026 · 14 citations
