IndiGEC: Multilingual Grammar Error Correction for Low-Resource Indian Languages
Ujjwal Sharma, Pushpak Bhattacharyya
Abstract
Grammatical Error Correction (GEC) for lowresource Indic languages faces significant challenges due to the scarcity of annotated data. In this work, we introduce the Mask-Translate&Fill (MTF) framework, a novel approach for generating high-quality synthetic data for GEC using only monolingual corpora. MTF leverages a machine translation system and a pretrained masked language model to introduce synthetic errors and tries to mimic errors made by second-language learners. Our experimental results on English, Hindi, Bengali, Marathi, and Tamil demonstrate that MTF consistently outperforms other monolingual synthetic data generation methods and achieves performance comparable to the Translation Language Modeling (TLM)-based approach, which uses a bilingual corpus, in both independent and multilingual settings. Under multilingual training, MTF yields significant improvements across Indic languages, with particularly notable gains in Bengali and Tamil, achieving +1.6 and +3.14 GLEU over the TLMbased method, respectively. To support further research, we also introduce the IndiGEC Corpus, a high-quality, human-written, manually validated GEC dataset for these four Indic languages, comprising over 8,000 sentence pairs with separate development and test splits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 486ebb50-2205-4017-9ab7-401d6d64297cRelated papers
- Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error DetectionGaetan Latouche, Marc-André Carbonneau, Benjamin SwansonEMNLP 2024 · 1 citation
- Pretraining Language Models Using TranslationeseMeet Doshi, Raj Dabre, Pushpak BhattacharyyaEMNLP 2024
- MaskGEC: Improving Neural Grammatical Error Correction via Dynamic MaskingZewei Zhao, Houfeng WangAAAI 2020 · 71 citations
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic LanguagesGuduru Manoj, Neel Prabhanjan Rachamalla, Ashish Kulkarni, Gautam Rajeev et al.AAAI 2026 · 2 citations
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan et al.ACL 2023 · 8 citations
