Targeted Syntactic Evaluation for Grammatical Error Correction
Aomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama, Mamoru Komachi
摘要
Language learners encounter a wide range of grammar items across the beginner, intermediate, and advanced levels. To develop grammatical error correction (GEC) models effectively, it is crucial to identify which grammar items are easier or more challenging for models to correct. However, conventional benchmarks based on learner-produced texts are insufficient for conducting detailed evaluations of GEC model performance across a wide range of grammar items due to biases in their distribution. To address this issue, we propose a new evaluation paradigm that assesses GEC models using minimal pairs of ungrammatical and grammatical sentences for each grammar item. As the first benchmark within this paradigm, we introduce the CEFR-based Targeted Syntactic Evaluation Dataset for Grammatical Error Correction (CTSEG), which complements existing English benchmarks by enabling fine-grained analyses previously unattainable with conventional datasets. Using CTSEG, we evaluate three mainstream types of English GEC models: sequence-to-sequence models, sequence tagging models, and prompt-based models. The results indicate that while current models perform well on beginner-level grammar items, their performance deteriorates substantially for intermediate and advanced items.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Towards standardizing Korean Grammatical Error Correction: Datasets and AnnotationSoyoung Yoon, Sungjoon Park, Gyuwan Kim, Junhee Cho 等ACL 2023 · 被引用 6 次
- Enhancing Grammatical Error Correction Systems with ExplanationsYuejiao Fei, Leyang Cui, Sen Yang, Wai Lam 等ACL 2023 · 被引用 13 次
- Sequence-to-Action: Grammatical Error Correction with Action Guided Sequence GenerationJiquan Li, Junliang Guo, Yongxin Zhu, Xin Sheng 等AAAI 2022 · 被引用 29 次
- RobustGEC: Robust Grammatical Error Correction Against Subtle Context PerturbationYue Zhang, Leyang Cui, Enbo Zhao, Wei Bi 等EMNLP 2023 · 被引用 1 次
- TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language ModelsJie He, Bo Peng, Yi Liao, Qun Liu 等ACL 2021
