Edit-Aware Reward Modeling for Chinese Grammatical Error Correction
Yilin Li, Xiaojun Wan
Abstract
While large language models have achieved remarkable success in various natural language processing tasks, their potential in grammatical error correction remains underexplored. Recent work has applied reinforcement learning with rule-based rewards to CGEC, but these approaches rely on coarse-grained binary signals (exact match or not) that fail to capture fine-grained quality distinctions among correction candidates. In this paper, we pro-pose Edit-Aware Reward Model (EARM) , a novel reward modeling framework that explicitly incorporates edit-awareness into preference learning for CGEC. EARM introduces a dual-granularity training objective that jointly optimizes sentence-level and token-level weighted Bradley-Terry ranking losses, where edit tokens receive higher importance weights. When integrated with GRPO, our approach achieves 61.29/63.08 F 0 . 5 on FCGEC/NaCGEC (single output), and 65.04/64.59 with best-of-16 reranking, surpassing previous best by 5.41 and 1.80 points. Extensive experiments demonstrate that learned edit-aware rewards significantly outperform rule-based alternatives for CGEC preference optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53e84927-bb6d-4d33-b543-34362bce363fBuilds on7
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented ParserYue Zhang, Bo Zhang, Zhenghua Li, Zuyi Bao et al.EMNLP 2022 · 35 citations
Related papers
- CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware RewardsWei Tian, Yuhao Zhou, Man LanACL 2026
- DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language ModelsJinxiang Xie, Yilin Li, Xunjian Yin, Xiaojun WanAAAI 2025 · 2 citations
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 9 citations
- Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error OvercorrectionTaehee Park, Heejin Do, Gary LeeEMNLP 2025 · 1 citation
- ALRMR-GEC: Adjusting Learning Rate Based on Memory Rate to Optimize the Edit Scorer for Grammatical Error CorrectionZhixiao Wu, Yao Lu, Jie Wen, Guangming LuAAAI 2025
