CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction
Jingheng Ye, Zishan Xu, Yinghui Li, Linlin Song, Qingyu Zhou, Hai-Tao Zheng, Ying Shen, Wenhao Jiang, Hong-Gee Kim, Ruitong Liu, Xin Su, Zifei Shan
Abstract
The paper focuses on the interpretability of Grammatical Error Correction (GEC) evaluation metrics, which received little attention in previous studies. To bridge the gap, we introduce CLEME2.0, a reference-based metric describing four fundamental aspects of GEC systems: hit-correction, wrong-correction, under-correction, and over-correction. They collectively contribute to exposing critical qualities and locating drawbacks of GEC systems. Evaluating systems by combining these aspects also leads to superior human consistency over other reference-based and reference-less metrics. Extensive experiments on two human judgment datasets and six reference datasets demonstrate the effectiveness and robustness of our method, achieving a new state-of-the-art result. Our codes are released at https://github.com/THUKElab/CLEME.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ee7b007-0486-4ef7-8f5e-8a7b36874565Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingTianyu Yu, Chengyue Jiang, Chao Lou, Shen Huang et al.AAAI 2024 · 30 citations
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye et al.ACL 2026 · 24 citations
- Contrastive Learning with Hard Negative Entities for Entity Set ExpansionYinghui Li, Yangning Li, Yuxin He, Tianyu Yu et al.SIGIR 2022 · 22 citations
- MESED: A Multi-Modal Entity Set Expansion Dataset with Fine-Grained Semantic Classes and Hard Negative EntitiesYangning Li, Tingwei Lu, Hai-Tao Zheng, Yinghui Li et al.AAAI 2024 · 21 citations
Related papers
- CLEME: Debiasing Multi-reference Evaluation for Grammatical Error CorrectionJingheng Ye, Yinghui Li, Qingyu Zhou, Yangning Li et al.EMNLP 2023 · 5 citations
- EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error CorrectionJingheng Ye, Shang Qin, Yinghui Li, Xuxin Cheng et al.AAAI 2025 · 3 citations
- DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language ModelsJinxiang Xie, Yilin Li, Xunjian Yin, Xiaojun WanAAAI 2025 · 2 citations
- Interpretability for Language Learners Using Example-Based Grammatical Error CorrectionMasahiro Kaneko, Sho Takase, Ayana Niwa, Naoaki OkazakiACL 2022
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 11 citations
