Efficient Grammatical Error Correction Via Multi-Task Training and Optimized Training Schedule
Andrey Bout, Alexander Podolskiy, Sergey I. Nikolenko, Irina Piontkovskaya
Abstract
Progress in neural grammatical error correction (GEC) is hindered by the lack of annotated training data. Sufficient amounts of high-quality manually annotated data are not available, so recent research has relied on generating synthetic data, pretraining on it, and then fine-tuning on real datasets; performance gains have been achieved either by ensembling or by using huge pretrained models such as XXL-T5 as the backbone. In this work, we explore an orthogonal direction: how to use available data more efficiently. First, we propose auxiliary tasks that exploit the alignment between the original and corrected sentences, such as predicting a sequence of corrections. We formulate each task as a sequence-to-sequence problem and perform multi-task training. Second, we discover that the order of datasets used for training and even individual instances within a dataset may have important effects on the final performance, so we set out to find the best training schedule. Together, these two ideas lead to significant improvements, producing results that improve state of the art with much smaller models; in particular, we outperform the best models based on T5-XXL (11B parameters) with a BART-based model (400M parameters).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 9 citations
- Multi-pass Decoding for Grammatical Error CorrectionXiaoying Wang, Lingling Mu, Jingyi Zhang, Hongfei XuEMNLP 2024 · 2 citations
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented ParserYue Zhang, Bo Zhang, Zhenghua Li, Zuyi Bao et al.EMNLP 2022 · 35 citations
- Ensembling and Knowledge Distilling of Large Sequence Taggers for Grammatical Error CorrectionMaksym Tarnavskyi, Artem N. Chernodub, Kostiantyn OmelianchukACL 2022 · 28 citations
Related papers
- Byte-Level Grammatical Error Correction Using Synthetic and Curated CorporaSvanhvít Lilja Ingólfsdóttir, Petur Orri Ragnarsson, Haukur Páll Jónsson, Haukur Barri Símonarson et al.ACL 2023 · 4 citations
- Multi-Class Grammatical Error Detection for Correction: A Tale of Two SystemsZheng Yuan, Shiva Taslimipoor, Christopher Davis, Christopher BryantEMNLP 2021 · 28 citations
- GEC-DePenD: Non-Autoregressive Grammatical Error Correction with Decoupled Permutation and DecodingKonstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko et al.ACL 2023 · 4 citations
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen et al.EMNLP 2021 · 176 citations
- On Losses for Modern Language ModelsStephane Aroca-Ouellette, Frank RudziczEMNLP 2020 · 2 citations
