Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization
Salman Elgamal, Ossama Obeid, Mhd Tameem Kabbani, Go Inoue, Nizar Habash
Abstract
The widespread absence of diacritical marks in Arabic text poses a significant challenge for Arabic natural language processing (NLP). This paper explores instances of naturally occurring diacritics, referred to as "diacritics in the wild," to unveil patterns and latent information across six diverse genres: news articles, novels, children's books, poetry, political documents, and ChatGPT outputs. We present a new annotated dataset that maps realworld partially diacritized words to their maximal full diacritization in context. Additionally, we propose extensions to the analyze-anddisambiguate approach in Arabic NLP to leverage these diacritics, resulting in notable improvements. Our contributions encompass a thorough analysis, valuable datasets, and an extended diacritization algorithm. We release our code and datasets as open source.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c40b418e-f2b2-4f91-a3fb-9a0a167db13eBuilds on3
- Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationBashar Alhafni, Go Inoue, Christian Khairallah, Nizar HabashEMNLP 2023 · 10 citations
- Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological TaggingNasser Zalmout, Nizar HabashACL 2020 · 3 citations
- A Multitask Learning Approach for Diacritic RestorationSawsan Alqahtani, Ajay Mishra, Mona T. DiabACL 2020 · 2 citations
Related papers
- Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art ModelsAbubakr Mohamed, Hamdy MubarakEMNLP 2025
- ALDi: Quantifying the Arabic Level of Dialectness of TextAmr Keleg, Sharon Goldwater, Walid MagdyEMNLP 2023 · 4 citations
- To Distill or Not to Distill? On the Robustness of Robust Knowledge DistillationAbdul Waheed, Karima Kadaoui, Muhammad Abdul-MageedACL 2024
- Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah et al.EMNLP 2024 · 5 citations
- Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsAbdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi et al.ACL 2026
