Revisiting Grammatical Error Correction Evaluation and Beyond
Peiyuan Gong, Xuebo Liu, Heyan Huang, Min Zhang
Abstract
Pretraining-based (PT-based) automatic evaluation metrics (e.g., BERTScore and BARTScore) have been widely used in several sentence generation tasks (e.g., machine translation and text summarization) due to their better correlation with human judgments over traditional overlapbased methods. Although PT-based methods have become the de facto standard for training grammatical error correction (GEC) systems, GEC evaluation still does not benefit from pretrained knowledge. This paper takes the first step towards understanding and improving GEC evaluation with pretraining. We first find that arbitrarily applying PT-based metrics to GEC evaluation brings unsatisfactory correlation results because of the excessive attention to inessential systems outputs (e.g., unchanged parts). To alleviate the limitation, we propose a novel GEC evaluation metric to achieve the best of both worlds, namely PT-M 2 , which only uses PT-based metrics to score those corrected parts. Experimental results on the CoNLL14 evaluation task show that PT-M 2 significantly outperforms existing methods, achieving a new state-of-the-art result of 0.949 Pearson correlation. Further analysis reveals that PT-M 2 is robust to evaluate competitive GEC systems. Source code and scripts are freely available at https://github.com/pygongnlp/PT-M2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73459e5f-ab26-4479-a4eb-1fca609c2d6aCited by top-tier papers6
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong et al.ACL 2023 · 21 citations
- CLEME: Debiasing Multi-reference Evaluation for Grammatical Error CorrectionJingheng Ye, Yinghui Li, Qingyu Zhou, Yangning Li et al.EMNLP 2023 · 5 citations
- Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error OvercorrectionTaehee Park, Heejin Do, Gary LeeEMNLP 2025 · 1 citation
- ALRMR-GEC: Adjusting Learning Rate Based on Memory Rate to Optimize the Edit Scorer for Grammatical Error CorrectionZhixiao Wu, Yao Lu, Jie Wen, Guangming LuAAAI 2025
- CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error CorrectionJingheng Ye, Zishan Xu, Yinghui Li, Linlin Song et al.ACL 2025
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation EvaluationWei Zhao, Goran Glavas, Maxime Peyrard, Yang Gao et al.ACL 2020 · 53 citations
- ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence GenerationBei Li, Quan Du, Tao Zhou, Yi Jing et al.ACL 2022 · 43 citations
Related papers
- Toward Human-Like Evaluation for Natural Language Generation with Error AnalysisQingyu Lu, Liang Ding, Liping Xie, Kanjian Zhang et al.ACL 2023 · 10 citations
- Improving Image Captioning Evaluation by Considering Inter References VarianceYanzhi Yi, Hangyu Deng, Jinglu HuACL 2020 · 44 citations
- BLEURT Has Universal Translations: An Analysis of Automatic Metrics by Minimum Risk TrainingYiming Yan, Tao Wang, Chengqi Zhao, Shujian Huang et al.ACL 2023 · 7 citations
- BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text GenerationTianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing HuangEMNLP 2022 · 22 citations
- System Combination via Quality Estimation for Grammatical Error CorrectionMuhammad Reza Qorib, Hwee Tou NgEMNLP 2023 · 8 citations
