ScholarGEC: Enhancing Controllability of Large Language Model for Chinese Academic Grammatical Error Correction
Zixiao Kong, Xianquan Wang, Shuanghong Shen, Keyu Zhu, Huibo Xu, Yu Su
Abstract
Large language models (LLMs) have demonstrated exceptional error detection capabilities and can correct sentences with high fluency in grammatical error correction (GEC) tasks. However, when correcting Chinese academic papers, LLMs face significant challenges of over-correction. To delve deeper into this issue, we explore the underlying reasons. On one hand, each discipline has its unique vocabulary and expressions, and LLMs have insufficient and incomplete understanding of domain-specific sentences. On the other hand, the controllability of generative LLMs in GEC tasks is inherently poor, and the traditional sequence-to-sequence (Seq2Seq) correction structure exacerbates this issue. Considering the two aforementioned factors, we propose a new error correction framework for Chinese academic GEC tasks using LLMs, named ScholarGEC. To improve LLMs' understanding of domain-specific knowledge, we construct appropriate disciplinary knowledge prefixes for sentences and use this domain-specific knowledge data to fine-tune the LLM. To enhance the controllability of LLMs, we replace the traditional Seq2Seq structure with a Detection-Correction separated structure. We also introduce a special token during the process to improve the model's error detection stability. Additionally, we incorporate iterative self-reflection to enhance the stability of the generation, in the three parts of LLM generation. Extensive experiments demonstrate the effectiveness and robustness of our framework on a Chinese GEC dataset composed of academic papers, and further analysis reveals the capabilities of our framework in enhancing LLM performance in general GEC tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9bbbefa-3aec-476e-9996-b87b37ec0710Cited by top-tier papers3
- CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware RewardsWei Tian, Yuhao Zhou, Man LanACL 2026
- ContrastKV: Robust KV Cache Eviction via Contrastive Signal Fusion for Multi-Query GeneralizationXingchi Chen, Peiyuan Zong, Ziqiang Gao, Qing Li et al.ACL 2026
- MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context InferenceKunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang et al.ACL 2025
Builds on5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong et al.ACL 2023 · 21 citations
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 9 citations
- Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language ModelsRui Li, Qi Liu, Liyang He, Zheng Zhang et al.EMNLP 2024 · 4 citations
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
Related papers
- Intuitive Thinking: Expanding Large Language Models' Thinking for Rapid Decision-Making on Candidate Corrections in Chinese Grammar Error CorrectionLintao Long, Ruizhang Huang, Ruina Bai, Yongbin Qin et al.AAAI 2026
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao et al.AAAI 2025 · 19 citations
- Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error OvercorrectionTaehee Park, Heejin Do, Gary LeeEMNLP 2025 · 1 citation
- CL²GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error CorrectionShang Qin, Jingheng Ye, Yinghui Li, Hai-Tao Zheng et al.ACL 2026
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
