Automatic Correction of Writing Anomalies in Hausa Texts
Ahmad Mustapha Wali, Sergiu Nisioi
Abstract
Hausa texts are often characterized by writing anomalies, such as incorrect character substitutions and spacing errors, which sometimes hinder natural language processing (NLP) applications. This paper presents an approach to automatically correct anomalies by finetuning transformer-based models. Using a corpus gathered from several public sources, we create a large-scale parallel dataset of over 400,000 noisy-clean Hausa sentence pairs by introducing synthetically generated noise to mimic realistic writing errors. In addition, we finetune several multilingual and African language models, including M2M100, AfriTeVA, NCAIR1/N-ATLaS, UBC-NLP/cheetah-base, and other variants of BART and T5 for this correction task. Our experimental results demonstrate that models such as M2M100 achieve state-of-the-art results despite their smaller size and distinct pretraining, and that correcting errors can have a significant impact in improving downstream tasks such as text classification, machine translation, question answering, and LLM prompting in general. This research provides a methodology, a publicly available dataset, and a comparison of models to improve Hausa text quality, thereby advancing NLP capabilities for the language and offering transferable insights for other low-resource languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75034539-7f37-439e-9d5c-e891c4ffa530Builds on3
- MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African languagesCheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba O. Alabi et al.ACL 2023 · 12 citations
- Cheetah: Natural Language Generation for 517 African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-MageedACL 2024 · 1 citation
- AFRIDOC-MT: Document-level MT Corpus for African LanguagesJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet et al.EMNLP 2025
Related papers
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
- INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African LanguagesHao Yu, Jesujoba Oluwadara Alabi, Andiswa Bukula, Jian Yun Zhuang et al.ACL 2025 · 8 citations
- LangMark: A Multilingual Dataset for Automatic Post-EditingDiego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin et al.ACL 2025
- Error Analysis of Multilingual Language Models in Machine Translation: A Case Study of English-Amharic TranslationHizkiel Alemayehu, Hamada M. Zahera, Axel-Cyrille Ngonga NgomoEMNLP 2024 · 1 citation
- ZINA: Multimodal Fine-grained Hallucination Detection and EditingYuiga Wada, Kazuki Matsuda, Komei Sugiura, Graham NeubigCVPR 2026 · 5 citations
