CLEME: Debiasing Multi-reference Evaluation for Grammatical Error Correction
Jingheng Ye, Yinghui Li, Qingyu Zhou, Yangning Li, Shirong Ma, Hai-Tao Zheng, Ying Shen
Abstract
Evaluating the performance of Grammatical Error Correction (GEC) systems is a challenging task due to its subjectivity. Designing an evaluation metric that is as objective as possible is crucial to the development of GEC task. However, mainstream evaluation metrics, i.e., reference-based metrics, introduce bias into the multi-reference evaluation by extracting edits without considering the presence of multiple references. To overcome this issue, we propose Chunk-LE Multi-reference Evaluation (CLEME), designed to evaluate GEC systems in the multi-reference evaluation setting. CLEME builds chunk sequences with consistent boundaries for the source, the hypothesis and references, thus eliminating the bias caused by inconsistent edit boundaries. Furthermore, we observe the consistent boundary could also act as the boundary of grammatical errors, based on which the F0.5 score is then computed following the correction independence assumption. We conduct experiments on six English reference sets based on the CoNLL-2014 shared task. Extensive experiments and detailed analyses demonstrate the correctness of our discovery and the effectiveness of CLEME. Further analysis reveals that CLEME is robust to evaluate GEC systems across reference sets with varying numbers of references and annotation styles. All the source codes of CLEME are released at https://github.com/THUKElab/CLEME.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0f41d53-1d9e-4717-b78e-c95daa063600Cited by top-tier papers4
- EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error CorrectionJingheng Ye, Shang Qin, Yinghui Li, Xuxin Cheng et al.AAAI 2025 · 3 citations
- CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error CorrectionJingheng Ye, Zishan Xu, Yinghui Li, Linlin Song et al.ACL 2025
- Toward Robust Evaluation for Multilingual Grammatical Error Correction: Can Large Language Models Replace Human References?Alla Rozovskaya, Dan RothACL 2026
- JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error CorrectionYuhao Zhan, Yuqing Zhang, Jing Yuan, Qixiang Ma et al.AAAI 2026
Builds on3
- On the Limitations of Reference-Free Evaluations of Generated TextDaniel Deutsch, Rotem Dror, Dan RothEMNLP 2022 · 23 citations
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 11 citations
- Interpretability for Language Learners Using Example-Based Grammatical Error CorrectionMasahiro Kaneko, Sho Takase, Ayana Niwa, Naoaki OkazakiACL 2022
Related papers
- System Combination via Quality Estimation for Grammatical Error CorrectionMuhammad Reza Qorib, Hwee Tou NgEMNLP 2023 · 8 citations
- DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language ModelsJinxiang Xie, Yilin Li, Xunjian Yin, Xiaojun WanAAAI 2025 · 2 citations
- Grammatical Error Correction in Low Error Density Domains: A New Benchmark and AnalysesSimon Flachs, Ophélie Lacroix, Helen Yannakoudakis, Marek Rei et al.EMNLP 2020
- Targeted Syntactic Evaluation for Grammatical Error CorrectionAomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama et al.ACL 2025
- RobustGEC: Robust Grammatical Error Correction Against Subtle Context PerturbationYue Zhang, Leyang Cui, Enbo Zhao, Wei Bi et al.EMNLP 2023 · 1 citation
