C-LLM: Learn to Check Chinese Spelling Errors Character by Character
Kunting Li, Yong Hu, Liang He, Fandong Meng, Jie Zhou
摘要
Chinese Spell Checking (CSC) aims to detect and correct spelling errors in sentences. Despite Large Language Models (LLMs) exhibit robust capabilities and are widely applied in various tasks, their performance on CSC is often unsatisfactory. We find that LLMs fail to meet the Chinese character-level constraints of the CSC task, namely equal length and phonetic similarity, leading to a performance bottleneck. Further analysis reveals that this issue stems from the granularity of tokenization, as current mixed character-word tokenization struggles to satisfy these characterlevel constraints. To address this issue, we propose C-LLM, a Large Language Modelbased Chinese Spell Checking method that learns to check errors Character by Character. Character-level tokenization enables the model to learn character-level alignment, effectively mitigating issues related to character-level constraints. Furthermore, CSC is simplified to replication-dominated and substitutionsupplemented tasks. Experiments on two CSC benchmarks demonstrate that C-LLM achieves an average improvement of 10% over existing methods. Specifically, it shows a 2.1% improvement in general scenarios and a significant 12% improvement in vertical domain scenarios, establishing state-of-the-art performance. The source code can be accessed at https://github.com/ktlKTL/C-LLM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated RewardsZhiming Lin, Kai Zhao, Sophie Zhang, Peilai Yu 等AAAI 2026 · 被引用 11 次
- Enhancing Character-Level Understanding in LLMs through Token Internal Structure LearningZhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu 等ACL 2025 · 被引用 7 次
- Mixture of Small and Large Models for Chinese Spelling CheckZiheng Qiao, Houquan Zhou, Zhenghua LiACL 2025 · 被引用 4 次
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li 等EMNLP 2024 · 被引用 2 次
- CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware RewardsWei Tian, Yuhao Zhou, Man LanACL 2026
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 被引用 204 次
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling CheckXingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang 等ACL 2020 · 被引用 139 次
相关 Paper
- ARM: An Alignment-and-Replacement Module for Chinese Spelling Check Based on LLMsChangchun Liu, Kai Zhang, Junzhe Jiang, Zirui Liu 等EMNLP 2024 · 被引用 3 次
- PLOME: Pre-training with Misspelled Knowledge for Chinese Spelling CorrectionShulin Liu, Tao Yang, Tianchi Yue, Feng Zhang 等ACL 2021
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 被引用 8 次
- PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling CheckLi Huang, Junjie Li, Weiwei Jiang, Zhiyu Zhang 等ACL 2021
- Chinese Spelling Correction as Rephrasing Language ModelLinfeng Liu, Hongqiu Wu, Hai ZhaoAAAI 2024 · 被引用 36 次
