NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, Muhammad Abdul-Mageed
摘要
How are you, Mostafa? Aren't you going to eat? Mostafa: Sorry, I was talking to the wife. This street food makes me nervous, right? Ahmed: Man, don't worry, this place is clean. Besides, this is Alexandrian koshary with yellow lentils, different from what you're used to. Mostafa: May God keep us safe. A man has to Figure 1: Our proposed framework enhances text data augmentation for low-resource local communities through a multi-stage pipeline. First, it (a) generates educational data using machine translation. Next, it (b) creates diverse, culturally-aware texts, such as stories and conversations, by simulating scenarios with local personas through controlled synthetic data generation. Finally, it (c) enriches the model with local knowledge by retrieving and parsing culturally specific web content. This entire process enables controlled text generation and retrievalaugmented pre-training, ensuring the cultural and value alignment of large language models for Arabic dialects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective ResponsesChongyuan Dai, Yaling Shen, Zihan Gao, Jia Li 等ACL 2026 · 被引用 4 次
- Cultural Benchmarking of LLMs in Standard and Dialectal Arabic DialoguesMuhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi 等ACL 2026
- Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsAbdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi 等ACL 2026
- Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsFakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany 等ACL 2025
它引用的顶会 Paper8
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram 等NeurIPS 2024 · 被引用 101 次
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal 等EMNLP 2024 · 被引用 53 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 被引用 27 次
- JASMINE: Arabic GPT Models for Few-Shot LearningEl Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Alcides Alcoba Inciarte 等EMNLP 2023 · 被引用 13 次
相关 Paper
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing 等ACL 2022 · 被引用 114 次
- DALE: Generative Data Augmentation for Low-Resource Legal NLPSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. 等EMNLP 2023 · 被引用 10 次
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun 等EMNLP 2021 · 被引用 22 次
- Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity RecognitionZiyan Li, Jianfei Yu, Jia Yang, Wenya Wang 等ACM MM 2024 · 被引用 13 次
- Having Beer after Prayer? Measuring Cultural Bias in Large Language ModelsTarek Naous, Michael J. Ryan, Alan Ritter, Wei XuACL 2024
