LM-Critic: Language Models for Unsupervised Grammatical Error Correction
Michihiro Yasunaga, Jure Leskovec, Percy Liang
Abstract
Training a model for grammatical error correction (GEC) requires a set of labeled ungrammatical / grammatical sentence pairs, but manually annotating such pairs can be expensive. Recently, the Break-It-Fix-It (BIFI) framework has demonstrated strong results on learning to repair a broken program without any labeled examples, but this relies on a perfect critic (e.g., a compiler) that returns whether an example is valid or not, which does not exist for the GEC task. In this work, we show how to leverage a pretrained language model (LM) in defining an LM-Critic, which judges a sentence to be grammatical if the LM assigns it a higher probability than its local perturbations. We apply this LM-Critic and BIFI along with a large set of unlabeled sentences to bootstrap realistic ungrammatical/grammatical pairs for training a corrector. We evaluate our approach on GEC datasets across multiple domains (CoNLL-2014, BEA-2019, GMEG-wiki and GMEG-yahoo) and show that it outperforms existing methods in both the unsupervised setting (+7.7 F 0.5 ) and the supervised setting (+0.5 F 0.5 ).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93441b2a-4f4f-4df9-a438-d8fd8dbe28cfCited by top-tier papers13
- SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented ParserYue Zhang, Bo Zhang, Zhenghua Li, Zuyi Bao et al.EMNLP 2022 · 35 citations
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong et al.ACL 2023 · 21 citations
- Converge to the Truth: Factual Error Correction via Iterative Constrained EditingJiangjie Chen, Rui Xu, Wenxuan Zeng, Changzhi Sun et al.AAAI 2023 · 13 citations
- Improved grammatical error correction by ranking elementary editsAlexey SorokinEMNLP 2022 · 10 citations
- Unsupervised Grammatical Error Correction Rivaling Supervised MethodsHannan Cao, Liping Yuan, Yuchen Zhang, Hwee Tou NgEMNLP 2023 · 4 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Break-It-Fix-It: Unsupervised Learning for Program RepairMichihiro Yasunaga, Percy LiangICML 2021 · 128 citations
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 92 citations
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 25 citations
Related papers
- JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error CorrectionYuhao Zhan, Yuqing Zhang, Jing Yuan, Qixiang Ma et al.AAAI 2026
- Improving Grammatical Error Correction Models with Purpose-Built Adversarial ExamplesLihao Wang, Xiaoqing ZhengEMNLP 2020 · 21 citations
- Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationBashar Alhafni, Go Inoue, Christian Khairallah, Nizar HabashEMNLP 2023 · 10 citations
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 9 citations
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 11 citations
