Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi, Younes Samih, Sarfraz Ahmad, Amr Keleg, Omar El Herraoui, Kareem Elzeky, Abed Alhakim Freihat, Mohamed Anwar, Zhuohan Xie
Abstract
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking the cultural nuances that naturally arise in dialogues. To address this gap, we introduce ArabCulture-Dialogue, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country's respective dialect, spanning 12 daily-life topics and 54 finegrained subtopics. We utilize the dataset to form three benchmarking tasks: (i) multiplechoice cultural reasoning, (ii) machine translation between MSA and dialects, and (iii) dialect-steering generation. Our experiments indicate that the performance gap between MSA and Arabic dialects still exists, whereby the models perform worse on all three tasks in the dialectal setup, compared to the MSA one. Recent years have seen substantial progress in Arabic NLP, with the emergence of Arabic-centric LLMs such as Jais (Sengupta et al., 2023), SILMA (SILMA-AI, 2024), and ALLaM (Bari et al., 2025) , alongside multilingual models that increasingly support Arabic. Evaluation benchmarks also have expanded accordingly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Commonsense Reasoning in Arab CultureAbdelrahman Boda Sadallah, Junior Cedric Tonga, Khalid Almubarak, Saeed Almheiri et al.ACL 2025 · 18 citations
- Peacock: A Family of Arabic Multimodal Large Language Models and BenchmarksFakhraddin Alwajih, El Moatez Billah Nagoudi, Gagan Bhatia, Abdelrahman Mohamed et al.ACL 2024 · 5 citations
- ALDi: Quantifying the Arabic Level of Dialectness of TextAmr Keleg, Sharon Goldwater, Walid MagdyEMNLP 2023 · 4 citations
- Challenges and Strategies in Cross-Cultural NLPDaniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux et al.ACL 2022
- NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local CommunitiesAbdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata et al.EMNLP 2025
Related papers
- Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsFakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany et al.ACL 2025
- Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsAbdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi et al.ACL 2026
- Having Beer after Prayer? Measuring Cultural Bias in Large Language ModelsTarek Naous, Michael J. Ryan, Alan Ritter, Wei XuACL 2024
- CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion UnderstandingTadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom II, Iqra Ameer et al.ACL 2025
- Seeing Culture: A Benchmark for Visual Reasoning and GroundingBurak Satar, Zhixin Ma, Patrick Amadeus Irawan, Wilfried A. Mulyawan et al.EMNLP 2025
