Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models
Abubakr Mohamed, Hamdy Mubarak
Abstract
Arabic diacritics, similar to short vowels in English, provide phonetic and grammatical information but are typically omitted in written Arabic, leading to ambiguity. Diacritization (aka diacritic restoration or vowelization) is essential for natural language processing. This paper advances Arabic diacritization through the following contributions: first, we propose a methodology to analyze and refine a large diacritized corpus to improve training data quality. Second, we introduce WikiNews-2024, a multi-reference evaluation methodology with an updated version of the standard benchmark "WikiNews-2014". In addition, we explore various model architectures and propose a BiLSTM-based model that achieves state-of-the-art results with 3.12% and 2.70% WER on WikiNews-2014 and WikiNews-2024, respectively. Moreover, we develop a model that preserves user-provided diacritics while maintaining accuracy. Lastly, we demonstrate that augmenting training data enhances performance in low-resource settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6db701ca-b105-4437-aaf0-195796d565a9Builds on1
Related papers
- Arabic Diacritics in the Wild: Exploiting Opportunities for Improved DiacritizationSalman Elgamal, Ossama Obeid, Mhd Tameem Kabbani, Go Inoue et al.ACL 2024 · 1 citation
- To Distill or Not to Distill? On the Robustness of Robust Knowledge DistillationAbdul Waheed, Karima Kadaoui, Muhammad Abdul-MageedACL 2024
- HintedBT: Augmenting Back-Translation with Quality and Transliteration HintsSahana Ramnath, Melvin Johnson, Abhirut Gupta, Aravindan RaghuveerEMNLP 2021 · 7 citations
- QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech CorpusHamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed AliACL 2021
- GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and RefinementYifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui et al.ACL 2025
