SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James V. Miranda, Jennifer Santoso, Elyanah Aco, Akhdan Fadhilah, Jonibek Mansurov, Joseph Marvin Imperial, Onno Kampman, Joel Ruben Antony Moniz, Muhammad Ravi Shulthan Habibi
Abstract
Southeast Asia (SEA) is a region characterized by rich linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, the performance of contemporary AI models for SEA languages is compromised by a significant lack of representation of texts, images, and auditory datasets from SEA. Evaluating models for SEA languages is challenging due to the scarcity of high-quality datasets, compounded by the predominance of English training data, which raises concerns regarding potential cultural misrepresentation. To address these challenges, we introduce SEACrowd, a collaborative initiative that consolidates a comprehensive resource hub 1 to bridge the resource gap by providing standardized corpora and benchmarks 2 in nearly 1,000 SEA languages across three modalities. We assess the performance of AI models on 36 indigenous languages across 13 tasks included in SEACrowd, offering valuable insights into the current AI landscape in SEA. Furthermore, we propose strategies to facilitate 1 https://seacrowd.github.io/seacrowd-catalogue/ 2 https://github.com/SEACrowd/seacrowd-datahub/ greater AI advancements, maximizing potential utility and resource equity for the future of AI in Southeast Asia.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7ad6ff3-450b-4dec-8541-a44e3f3b22e0Cited by top-tier papers11
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos et al.ICLR 2025 · 10 citations
- Kaleidoscope: In-language Exams for Massively Multilingual Vision EvaluationIsrafel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar et al.ICLR 2026 · 8 citations
- Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear ControllabilityXin Zhao, Zehui Jiang, Naoki YoshinagaACL 2025 · 3 citations
- LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian LanguagesAlham Fikri Aji, Trevor CohnEMNLP 2025 · 2 citations
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani et al.EMNLP 2022 · 46 citations
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetAshish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu SoricutEMNLP 2022 · 31 citations
Related papers
- Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast AsiaSamuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz, Tack Hwa Wong et al.ACL 2025
- SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng, Jann Railey Montalan et al.ACL 2026 · 2 citations
- SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast AsiaPengfei Yue, Xingran Zhao, Juntao Chen, Peng Hou et al.CVPR 2026 · 2 citations
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera et al.ACL 2026 · 5 citations
- Afri-MCQA: Multimodal Cultural Question Answering for African LanguagesAtnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime et al.ACL 2026 · 2 citations
