Rethinking Masked Language Modeling for Chinese Spelling Correction
Hongqiu Wu, Shaohua Zhang, Yuchen Zhang, Hai Zhao
Abstract
In this paper, we study Chinese Spelling Correction (CSC) as a joint decision made by two separate models: a language model and an error model. Through empirical analysis, we find that fine-tuning BERT tends to over-fit the error model while under-fit the language model, resulting in poor generalization to out-of-distribution error patterns. Given that BERT is the backbone of most CSC models, this phenomenon has a significant negative impact. To address this issue, we are releasing a multi-domain benchmark LEMON, with higher quality and diversity than existing benchmarks, to allow a comprehensive assessment of the open domain generalization of CSC models. Then, we demonstrate that a very simple strategy – randomly masking 20% non-error tokens from the input sequence during fine-tuning – is sufficient for learning a much better language model without sacrificing the error model. This technique can be applied to any model architecture and achieves new state-of-the-art results on SIGHAN, ECSpell, and LEMON.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8618ed72-ee24-4654-92da-87486e811fa6Cited by top-tier papers13
- Chinese Spelling Correction as Rephrasing Language ModelLinfeng Liu, Hongqiu Wu, Hai ZhaoAAAI 2024 · 36 citations
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated RewardsZhiming Lin, Kai Zhao, Sophie Zhang, Peilai Yu et al.AAAI 2026 · 11 citations
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng et al.EMNLP 2024 · 9 citations
- Towards Real-World Writing Assistance: A Chinese Character Checking Benchmark with Faked and Misspelled CharactersYinghui Li, Zishan Xu, Shaoshen Chen, Haojing Huang et al.ACL 2024 · 9 citations
- Enhancing Character-Level Understanding in LLMs through Token Internal Structure LearningZhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu et al.ACL 2025 · 7 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 204 citations
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling CheckXingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang et al.ACL 2020 · 139 citations
- MaskGEC: Improving Neural Grammatical Error Correction via Dynamic MaskingZewei Zhao, Houfeng WangAAAI 2020 · 71 citations
- Toward Adversarial Training on Contextualized Language RepresentationHongqiu Wu, Yongxiang Liu, Hanwen Shi, Hai Zhao et al.ICLR 2023 · 4 citations
Related papers
- A Training-free LLM-based Approach to General Chinese Character Error CorrectionHouquan Zhou, Bo Zhang, Zhenghua Li, Ming Yan et al.ACL 2025
- Mixture of Small and Large Models for Chinese Spelling CheckZiheng Qiao, Houquan Zhou, Zhenghua LiACL 2025 · 4 citations
- PLOME: Pre-training with Misspelled Knowledge for Chinese Spelling CorrectionShulin Liu, Tao Yang, Tianchi Yue, Feng Zhang et al.ACL 2021
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li et al.EMNLP 2024 · 2 citations
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 8 citations
