Disentangled Phonetic Representation for Chinese Spelling Correction
Zihong Liang, Xiaojun Quan, Qifan Wang
Abstract
Chinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts. Although efforts have been made to introduce phonetic information (Hanyu Pinyin) in this task, they typically merge phonetic representations with character representations, which tends to weaken the representation effect of normal texts. In this work, we propose to disentangle the two types of features to allow for direct interaction between textual and phonetic information. To learn useful phonetic representations, we introduce a pinyin-to-character objective to ask the model to predict the correct characters based solely on phonetic information, where a separation mask is imposed to disable attention from phonetic input to text. To avoid overfitting the phonetics, we further design a self-distillation module to ensure that semantic information plays a major role in the prediction. Extensive experiments on three CSC benchmarks demonstrate the superiority of our method in using phonetic information 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f7dac1-f737-400f-89e6-4b91d1ff09c3Cited by top-tier papers8
- Enhancing Character-Level Understanding in LLMs through Token Internal Structure LearningZhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu et al.ACL 2025 · 7 citations
- Mixture of Small and Large Models for Chinese Spelling CheckZiheng Qiao, Houquan Zhou, Zhenghua LiACL 2025 · 4 citations
- ARM: An Alignment-and-Replacement Module for Chinese Spelling Check Based on LLMsChangchun Liu, Kai Zhang, Junzhe Jiang, Zirui Liu et al.EMNLP 2024 · 3 citations
- Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text CorrectionXuan Feng, Tianlong Gu, Xiaoli Liu, Liang ChangAAAI 2025 · 2 citations
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li et al.EMNLP 2024 · 2 citations
Builds on4
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 204 citations
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling CheckXingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang et al.ACL 2020 · 139 citations
- PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling CheckLi Huang, Junjie Li, Weiwei Jiang, Zhiyu Zhang et al.ACL 2021
Related papers
- DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling CheckZiheng Qiao, Houquan Zhou, Yumeng Liu, Zhenghua Li et al.ACL 2025
- PLOME: Pre-training with Misspelled Knowledge for Chinese Spelling CorrectionShulin Liu, Tao Yang, Tianchi Yue, Feng Zhang et al.ACL 2021
- Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and GranularityJiahao Li, Quan Wang, Zhendong Mao, Junbo Guo et al.EMNLP 2022 · 19 citations
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 8 citations
- Chinese Spelling Correction as Rephrasing Language ModelLinfeng Liu, Hongqiu Wu, Hai ZhaoAAAI 2024 · 36 citations
