NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, Muhammad Abdul-Mageed
Abstract
How are you, Mostafa? Aren't you going to eat? Mostafa: Sorry, I was talking to the wife. This street food makes me nervous, right? Ahmed: Man, don't worry, this place is clean. Besides, this is Alexandrian koshary with yellow lentils, different from what you're used to. Mostafa: May God keep us safe. A man has to Figure 1: Our proposed framework enhances text data augmentation for low-resource local communities through a multi-stage pipeline. First, it (a) generates educational data using machine translation. Next, it (b) creates diverse, culturally-aware texts, such as stories and conversations, by simulating scenarios with local personas through controlled synthetic data generation. Finally, it (c) enriches the model with local knowledge by retrieving and parsing culturally specific web content. This entire process enables controlled text generation and retrievalaugmented pre-training, ensuring the cultural and value alignment of large language models for Arabic dialects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea1606a1-cb5f-489c-8f94-a4f490a48086Cited by top-tier papers4
- Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective ResponsesChongyuan Dai, Yaling Shen, Zihan Gao, Jia Li et al.ACL 2026 · 4 citations
- Cultural Benchmarking of LLMs in Standard and Dialectal Arabic DialoguesMuhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi et al.ACL 2026
- Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsAbdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi et al.ACL 2026
- Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsFakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany et al.ACL 2025
Builds on8
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram et al.NeurIPS 2024 · 101 citations
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal et al.EMNLP 2024 · 53 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 27 citations
- JASMINE: Arabic GPT Models for Few-Shot LearningEl Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Alcides Alcoba Inciarte et al.EMNLP 2023 · 13 citations
Related papers
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing et al.ACL 2022 · 114 citations
- DALE: Generative Data Augmentation for Low-Resource Legal NLPSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. et al.EMNLP 2023 · 10 citations
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun et al.EMNLP 2021 · 22 citations
- Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity RecognitionZiyan Li, Jianfei Yu, Jia Yang, Wenya Wang et al.ACM MM 2024 · 13 citations
- Having Beer after Prayer? Measuring Cultural Bias in Large Language ModelsTarek Naous, Michael J. Ryan, Alan Ritter, Wei XuACL 2024
