Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs
Abdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-Hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod
摘要
Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite this, machine translation (MT) systems often generalize poorly to dialectal input, limiting their utility for millions of speakers. We introduce Alexandria, a large-scale, community-driven, human-translated dataset designed to bridge this gap. Alexandria covers 13 Arab countries and 11 high-impact domains, including health, education, and agriculture. Unlike previous resources, Alexandria provides unprecedented granularity by associating contributions with city-of-origin metadata, capturing authentic local varieties beyond coarse regional labels. The dataset consists of parallel English-Dialectal Arabic multi-turn conversational scenarios annotated with speaker-addressee gender configurations, enabling the study of gender-conditioned variation in dialectal use. Comprising 107K total turns, Alexandria serves as both a training resource and as a rigorous benchmark for evaluating MT and Large Language Models (LLMs). Our automatic and human evaluation benchmarks the current capabilities of Arabic-aware LLMs in translating across diverse Arabic dialects and sub-dialects while exposing significant persistent challenges. The Alexandria dataset, the creation prompts, the translation and revision guidelines, and the evaluation code are publicly available in the following repository: https://github.com/UBC-NLP/Alexandria
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Toward Micro-Dialect Identification in Diaglossic and Code-Switched EnvironmentsMuhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim A. Elmadany, Lyle H. UngarEMNLP 2020
- NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local CommunitiesAbdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata 等EMNLP 2025
- Languages Still Left Behind: Toward a Better Multilingual Machine Translation BenchmarkChihiro Taguchi, Seng Mai, Keita Kurabe, Yusuke Sakai 等EMNLP 2025
相关 Paper
- Cultural Benchmarking of LLMs in Standard and Dialectal Arabic DialoguesMuhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi 等ACL 2026
- Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsFakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany 等ACL 2025
- Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah 等EMNLP 2024 · 被引用 5 次
- Dialectal Coverage And Generalization in Arabic Speech RecognitionAmirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan, Abdullah Alatir 等ACL 2025
- To Distill or Not to Distill? On the Robustness of Robust Knowledge DistillationAbdul Waheed, Karima Kadaoui, Muhammad Abdul-MageedACL 2024
