BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji
Abstract
People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition--an umbrella term for several NLP tasks--impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities in research efforts and proposed solutions, particularly for under-resourced languages, which often lack high-quality annotated datasets. In this paper, we present BRIGHTER--a collection of multi-labeled, emotion-annotated datasets in 28 different languages and across several domains. BRIGHTER primarily covers low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. We highlight the challenges related to the data collection and annotation processes, and then report experimental results for monolingual and crosslingual multi-label emotion identification, as well as emotion intensity recognition. We analyse the variability in performance across languages and text domains, both with and without the use of LLMs, and show that the BRIGHTER datasets represent a meaningful step towards addressing the gap in text-based emotion recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 773e0f05-a033-47b0-bf06-c9a162999fa6Cited by top-tier papers3
- AdaReasoner: Adaptive Reasoning Enables More Flexible ThinkingXiangqi Wang, Yue Huang, Yanbo Wang, Xiaonan Luo et al.NeurIPS 2025 · 21 citations
- Building Better: Avoiding Pitfalls in Developing Language Resources when Data is ScarceNedjma Ousidhoum, Meriem Beloucif, Saif M. MohammadACL 2025 · 5 citations
- DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment AnalysisLung-Hao Lee, Liang-Chih Yu, Natalia V. Loukachevitch, Ilseyar Alimova et al.ACL 2026 · 1 citation
Builds on8
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- AfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesShamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum et al.EMNLP 2023 · 33 citations
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen et al.ACL 2020 · 16 citations
- No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 LanguagesYoussef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov et al.EMNLP 2024 · 3 citations
- WorryWords: Norms of Anxiety Association for over 44k English WordsSaif MohammadEMNLP 2024 · 2 citations
Related papers
- MASIVE: Open-Ended Affective State Identification in English and SpanishNicholas Deas, Elsbeth Turcan, Iván Pérez Mejía, Kathleen R. McKeownEMNLP 2024 · 3 citations
- CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion UnderstandingTadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom II, Iqra Ameer et al.ACL 2025
- Multi-resolution Annotations for Emoji PredictionWeicheng Ma, Ruibo Liu, Lili Wang, Soroush VosoughiEMNLP 2020 · 6 citations
- Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language ModelsNicholas Deas, Iván Ernesto Pérez Mejía, Ellie Yang, Kathleen McKeownACL 2026
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
