Lune

ACL2022Top-tier venue

Ensembling and Knowledge Distilling of Large Sequence Taggers for Grammatical Error Correction

Maksym Tarnavskyi, Artem N. Chernodub, Kostiantyn Omelianchuk

2022Year
28Citations
11Top-tier citations

Abstract

In this paper, we investigate improvements to the GEC sequence tagging architecture with a focus on ensembling of recent cutting-edge Transformer-based encoders in Large configurations. We encourage ensembling models by majority votes on span-level edits because this approach is tolerant to the model architecture and vocabulary size. Our best ensemble achieves a new SOTA result with an F 0.5 score of 76.05 on BEA-2019 (test), even without pretraining on synthetic datasets. In addition, we perform knowledge distillation with a trained ensemble to generate new synthetic training datasets, "Troy-Blogs" and "Troy-1BW". Our best single sequence tagging model that is pretrained on the generated Troy-datasets in combination with the publicly available synthetic PIE dataset achieves a near-SOTA 1 result with an F 0.5 score of 73.21 on BEA-2019 (test). The code, datasets, and trained models are publicly available. 2 * This research was performed during Maksym Tarnavskyi's work on Ms.Sc. thesis at Ukrainian Catholic University (Tarnavskyi, 2021). 1 To the best of our knowledge, our best single model gives way only to much heavier T5 model (Rothe et al., 2021) .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 013c05e7-b8c8-4413-915c-4cdc1d76e1e6

Cited by top-tier papers11

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines