MASIVE: Open-Ended Affective State Identification in English and Spanish
Nicholas Deas, Elsbeth Turcan, Iván Pérez Mejía, Kathleen R. McKeown
摘要
In the field of emotion analysis, much NLP research focuses on identifying a limited number of discrete emotion categories, often applied across languages. These basic sets, however, are rarely designed with textual data in mind, and culture, language, and dialect can influence how particular emotions are interpreted. In this work, we broaden our scope to a practically unbounded set of affective states, which includes any terms that humans use to describe their experiences of feeling. We collect and publish MASIVE, a dataset of Reddit posts in English and Spanish containing over 1,000 unique affective states each. We then define the new problem of affective state identification for language generation models framed as a masked span prediction task. On this task, we find that smaller fine-tuned multilingual models outperform much larger LLMs, even on region-specific Spanish affective states. Additionally, we show that pre-training on MA-SIVE improves model performance on existing emotion benchmarks. Finally, through machine translation experiments, we find that native speaker-written data is vital to good performance on this task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related LanguagesVaidehi Patil, Partha P. Talukdar, Sunita SarawagiACL 2022 · 被引用 39 次
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen 等ACL 2020 · 被引用 16 次
- Evaluation of African American Language Bias in Natural Language GenerationNicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton 等EMNLP 2023 · 被引用 15 次
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 被引用 13 次
相关 Paper
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsZheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun 等ICML 2025
- Learning and Evaluating Emotion Lexicons for 91 LanguagesSven Buechel, Susanna Rücker, Udo HahnACL 2020 · 被引用 2 次
- BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle 等ACL 2025 · 被引用 81 次
- 3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short VideosVikram Gupta, Trisha Mittal, Puneet Mathur, Vaibhav Mishra 等CVPR 2022 · 被引用 14 次
- A Text-Based Recommender System that Leverages Explicit Affective State PreferencesTonmoy Hasan, Razvan C. BunescuEMNLP 2025 · 被引用 1 次
