A Multitask Learning Approach for Diacritic Restoration
Sawsan Alqahtani, Ajay Mishra, Mona T. Diab
Abstract
In many languages like Arabic, diacritics are used to specify pronunciations as well as meanings. Such diacritics are often omitted in written text, increasing the number of possible pronunciations and meanings for a word. This results in a more ambiguous text making computational processing on such text more difficult. Diacritic restoration is the task of restoring missing diacritics in the written text. Most state-of-the-art diacritic restoration models are built on character level information which helps generalize the model to unseen data, but presumably lose useful information at the word level. Thus, to compensate for this loss, we investigate the use of multi-task learning to jointly optimize diacritic restoration with related NLP problems namely word segmentation, part-of-speech tagging, and syntactic diacritization. We use Arabic as a case study since it has sufficient data resources for tasks that we consider in our joint modeling. Our joint models significantly outperform the baselines and are comparable to the state-of-the-art models that are more complex relying on morphological analyzers and/or a lot more data (e.g. dialectal data).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Arabic Diacritics in the Wild: Exploiting Opportunities for Improved DiacritizationSalman Elgamal, Ossama Obeid, Mhd Tameem Kabbani, Go Inoue et al.ACL 2024 · 1 citation
- Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art ModelsAbubakr Mohamed, Hamdy MubarakEMNLP 2025
Builds on1
Related papers
- To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource SettingsSarah R. Moeller, Ling Liu, Mans HuldenACL 2021
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 8 citations
- QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech CorpusHamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed AliACL 2021
- Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationBashar Alhafni, Go Inoue, Christian Khairallah, Nizar HabashEMNLP 2023 · 10 citations
- Getting The Most Out of Your Training Data: Exploring Unsupervised Tasks for Morphological InflectionAbhishek Purushothama, Adam Wiemerslage, Katharina von der WenseEMNLP 2024
