JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error Correction
Yuhao Zhan, Yuqing Zhang, Jing Yuan, Qixiang Ma, Zhiqi Yang, Yu Gu, Zemin Liu, Fei Wu
摘要
Existing Grammatical Error Correction (GEC) systems suffer from limited reference diversity, leading to underestimated evaluation and restricted model generalization. To address this issue, we introduce the Judge of Edit-Level Validity (JELV), an automated framework to validate correction edits from grammaticality, faithfulness, and fluency. Using our proposed human-annotated Pair-wise Edit-level Validity Dataset (PEVData) as benchmark, JELV offers two implementations: a multi-turn LLM-as-Judges pipeline achieving 90% agreement with human annotators, and a distilled DeBERTa classifier with 85% precision on valid edits. We then apply JELV to reclassify misjudged false positives in evaluation and derive a comprehensive evaluation metric by integrating false positive decoupling and fluency scoring, resulting in state-of-the-art correlation with human judgments. We also apply JELV to filter LLM-generated correction candidates, expanding the BEA19's single-reference dataset containing 38,692 source sentences. Retraining top GEC systems on this expanded dataset yields measurable performance gains. JELV provides a scalable solution for enhancing reference diversity and strengthening both evaluation and model generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun 等ICLR 2020 · 被引用 502 次
- LM-Critic: Language Models for Unsupervised Grammatical Error CorrectionMichihiro Yasunaga, Jure Leskovec, Percy LiangEMNLP 2021 · 被引用 29 次
- On the Limitations of Reference-Free Evaluations of Generated TextDaniel Deutsch, Rotem Dror, Dan RothEMNLP 2022 · 被引用 23 次
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 被引用 11 次
相关 Paper
- DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language ModelsJinxiang Xie, Yilin Li, Xunjian Yin, Xiaojun WanAAAI 2025 · 被引用 2 次
- Improved grammatical error correction by ranking elementary editsAlexey SorokinEMNLP 2022 · 被引用 10 次
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 被引用 9 次
- CLEME: Debiasing Multi-reference Evaluation for Grammatical Error CorrectionJingheng Ye, Yinghui Li, Qingyu Zhou, Yangning Li 等EMNLP 2023 · 被引用 5 次
- CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error CorrectionJingheng Ye, Zishan Xu, Yinghui Li, Linlin Song 等ACL 2025
