LM-Critic: Language Models for Unsupervised Grammatical Error Correction
Michihiro Yasunaga, Jure Leskovec, Percy Liang
摘要
Training a model for grammatical error correction (GEC) requires a set of labeled ungrammatical / grammatical sentence pairs, but manually annotating such pairs can be expensive. Recently, the Break-It-Fix-It (BIFI) framework has demonstrated strong results on learning to repair a broken program without any labeled examples, but this relies on a perfect critic (e.g., a compiler) that returns whether an example is valid or not, which does not exist for the GEC task. In this work, we show how to leverage a pretrained language model (LM) in defining an LM-Critic, which judges a sentence to be grammatical if the LM assigns it a higher probability than its local perturbations. We apply this LM-Critic and BIFI along with a large set of unlabeled sentences to bootstrap realistic ungrammatical/grammatical pairs for training a corrector. We evaluate our approach on GEC datasets across multiple domains (CoNLL-2014, BEA-2019, GMEG-wiki and GMEG-yahoo) and show that it outperforms existing methods in both the unsupervised setting (+7.7 F 0.5 ) and the supervised setting (+0.5 F 0.5 ).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented ParserYue Zhang, Bo Zhang, Zhenghua Li, Zuyi Bao 等EMNLP 2022 · 被引用 35 次
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong 等ACL 2023 · 被引用 21 次
- Converge to the Truth: Factual Error Correction via Iterative Constrained EditingJiangjie Chen, Rui Xu, Wenxuan Zeng, Changzhi Sun 等AAAI 2023 · 被引用 13 次
- Improved grammatical error correction by ranking elementary editsAlexey SorokinEMNLP 2022 · 被引用 10 次
- Unsupervised Grammatical Error Correction Rivaling Supervised MethodsHannan Cao, Liping Yuan, Yuchen Zhang, Hwee Tou NgEMNLP 2023 · 被引用 4 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Break-It-Fix-It: Unsupervised Learning for Program RepairMichihiro Yasunaga, Percy LiangICML 2021 · 被引用 128 次
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 被引用 92 次
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 被引用 25 次
相关 Paper
- JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error CorrectionYuhao Zhan, Yuqing Zhang, Jing Yuan, Qixiang Ma 等AAAI 2026
- Improving Grammatical Error Correction Models with Purpose-Built Adversarial ExamplesLihao Wang, Xiaoqing ZhengEMNLP 2020 · 被引用 21 次
- Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationBashar Alhafni, Go Inoue, Christian Khairallah, Nizar HabashEMNLP 2023 · 被引用 10 次
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 被引用 9 次
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 被引用 11 次
