Ensembling and Knowledge Distilling of Large Sequence Taggers for Grammatical Error Correction
Maksym Tarnavskyi, Artem N. Chernodub, Kostiantyn Omelianchuk
Abstract
In this paper, we investigate improvements to the GEC sequence tagging architecture with a focus on ensembling of recent cutting-edge Transformer-based encoders in Large configurations. We encourage ensembling models by majority votes on span-level edits because this approach is tolerant to the model architecture and vocabulary size. Our best ensemble achieves a new SOTA result with an F 0.5 score of 76.05 on BEA-2019 (test), even without pretraining on synthetic datasets. In addition, we perform knowledge distillation with a trained ensemble to generate new synthetic training datasets, "Troy-Blogs" and "Troy-1BW". Our best single sequence tagging model that is pretrained on the generated Troy-datasets in combination with the publicly available synthetic PIE dataset achieves a near-SOTA 1 result with an F 0.5 score of 73.21 on BEA-2019 (test). The code, datasets, and trained models are publicly available. 2 * This research was performed during Maksym Tarnavskyi's work on Ms.Sc. thesis at Ukrainian Catholic University (Tarnavskyi, 2021). 1 To the best of our knowledge, our best single model gives way only to much heavier T5 model (Rothe et al., 2021) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 013c05e7-b8c8-4413-915c-4cdc1d76e1e6Cited by top-tier papers11
- Enhancing Grammatical Error Correction Systems with ExplanationsYuejiao Fei, Leyang Cui, Sen Yang, Wai Lam et al.ACL 2023 · 13 citations
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 11 citations
- System Combination via Quality Estimation for Grammatical Error CorrectionMuhammad Reza Qorib, Hwee Tou NgEMNLP 2023 · 8 citations
- GEC-DePenD: Non-Autoregressive Grammatical Error Correction with Decoupled Permutation and DecodingKonstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko et al.ACL 2023 · 4 citations
- Multi-pass Decoding for Grammatical Error CorrectionXiaoying Wang, Lingling Mu, Jingyi Zhang, Hongfei XuEMNLP 2024 · 2 citations
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
Related papers
- Improved grammatical error correction by ranking elementary editsAlexey SorokinEMNLP 2022 · 10 citations
- Multi-Class Grammatical Error Detection for Correction: A Tale of Two SystemsZheng Yuan, Shiva Taslimipoor, Christopher Davis, Christopher BryantEMNLP 2021 · 28 citations
- Sequence-to-Action: Grammatical Error Correction with Action Guided Sequence GenerationJiquan Li, Junliang Guo, Yongxin Zhu, Xin Sheng et al.AAAI 2022 · 29 citations
- Generative Bias for Robust Visual Question AnsweringJae-Won Cho, Dong-Jin Kim, Hyeonggon Ryu, In So KweonCVPR 2023
- Efficient Grammatical Error Correction Via Multi-Task Training and Optimized Training ScheduleAndrey Bout, Alexander Podolskiy, Sergey I. Nikolenko, Irina PiontkovskayaEMNLP 2023 · 1 citation
