AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, Ibrahim Said Ahmad, Meriem Beloucif, Saif M. Mohammad, Sebastian Ruder, Oumaima Hourrane, Alípio Jorge
摘要
Africa is home to over 2,000 languages from more than six language families and has the highest linguistic diversity among all continents. These include 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial to enabling such research is the availability of high-quality annotated datasets. In this paper, we introduce AfriSenti, a sentiment analysis benchmark that contains a total of >110,000 tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yorùbá) from four language families. The tweets were annotated by native speakers and used in the AfriSenti-SemEval shared task 1 . We describe the data collection methodology, annotation process, and the challenges we dealt with when curating each dataset. We further report baseline experiments conducted on the different datasets and discuss their usefulness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle 等ACL 2025 · 被引用 81 次
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos 等ICLR 2025 · 被引用 10 次
- Voices Unheard: NLP Resources and Models for Yorùbá Regional DialectsOrevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen 等EMNLP 2024 · 被引用 1 次
- Culture Cartography: Mapping the Landscape of Cultural KnowledgeCaleb Ziems, William Barr Held, Jane Yu, Amir Goldberg 等EMNLP 2025
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson 等ACL 2024
它引用的顶会 Paper6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu 等ACL 2020 · 被引用 148 次
- MSCTD: A Multimodal Sentiment Chat Translation DatasetYunlong Liang, Fandong Meng, Jinan Xu, Yufeng Chen 等ACL 2022 · 被引用 27 次
- Discrete Opinion Tree Induction for Aspect-based Sentiment AnalysisChenhua Chen, Zhiyang Teng, Zhongqing Wang, Yue ZhangACL 2022
相关 Paper
- GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingMingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen 等EMNLP 2023
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadJesujoba Oluwadara Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich KlakowEMNLP 2025
- AfroLID: A Neural Language Identification Tool for African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba InciarteEMNLP 2022 · 被引用 13 次
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani 等EMNLP 2022 · 被引用 46 次
- Voice of a Continent: Mapping Africa's Speech Technology FrontierAbdelRahim A. Elmadany, Sang Yun Kwon, Hawau Olamide Toyin, Alcides Alcoba Inciarte 等EMNLP 2025
