A Training-free LLM-based Approach to General Chinese Character Error Correction
Houquan Zhou, Bo Zhang, Zhenghua Li, Ming Yan, Min Zhang
摘要
Chinese spelling correction (CSC) is a crucial task that aims to correct character errors in Chinese text. While conventional CSC focuses on character substitution errors caused by mistyping, two other common types of character errors, missing and redundant characters, have received less attention. These errors are often excluded from CSC datasets during the annotation process or ignored during evaluation, even when they have been annotated. This issue limits the practicality of the CSC task. To address this issue, we introduce the task of General Chinese Character Error Correction (C2EC), which focuses on all three types of character errors. We construct a high-quality C2EC benchmark by combining and manually verifying data from CCTC and Lemon datasets. We extend the training-free prompt-free CSC method to C2EC by using Levenshtein distance for handling length changes and leveraging an additional prompt-based large language model (LLM) to improve performance. Experiments show that our method enables a 14B-parameter LLM to be on par with models nearly 50 times larger on both conventional CSC and C2EC tasks, without any fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware RewardsWei Tian, Yuhao Zhou, Man LanACL 2026
- MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical DocumentsYijun Sheng, Shipeng Zhu, Ruijia Zuo, Na Nie 等CVPR 2026
它引用的顶会 Paper8
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 被引用 204 次
- Chinese Spelling Correction as Rephrasing Language ModelLinfeng Liu, Hongqiu Wu, Hai ZhaoAAAI 2024 · 被引用 36 次
- Rethinking Masked Language Modeling for Chinese Spelling CorrectionHongqiu Wu, Shaohua Zhang, Yuchen Zhang, Hai ZhaoACL 2023 · 被引用 21 次
- Disentangled Phonetic Representation for Chinese Spelling CorrectionZihong Liang, Xiaojun Quan, Qifan WangACL 2023 · 被引用 12 次
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng 等EMNLP 2024 · 被引用 9 次
相关 Paper
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li 等EMNLP 2024 · 被引用 2 次
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 被引用 8 次
- ARM: An Alignment-and-Replacement Module for Chinese Spelling Check Based on LLMsChangchun Liu, Kai Zhang, Junzhe Jiang, Zirui Liu 等EMNLP 2024 · 被引用 3 次
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated RewardsZhiming Lin, Kai Zhao, Sophie Zhang, Peilai Yu 等AAAI 2026 · 被引用 11 次
- PLOME: Pre-training with Misspelled Knowledge for Chinese Spelling CorrectionShulin Liu, Tao Yang, Tianchi Yue, Feng Zhang 等ACL 2021
