CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
Yong Hu, Fandong Meng, Jie Zhou
摘要
In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform. Compared with existing CSC datasets aimed at Chinese learners, CSCD-NS is ten times larger in scale and exhibits a distinct error distribution, with a significantly higher proportion of word-level errors. To further enhance the data resource, we propose a novel method that simulates the input process through an input method, generating large-scale and high-quality pseudo data that closely resembles the actual error distribution and outperforms existing methods. Moreover, we investigate the performance of various models in this scenario, including large language models (LLMs), such as ChatGPT. The result indicates that generative models underperform BERT-like classification models due to strict length and pronunciation constraints. The high prevalence of word-level errors also makes CSC for native speakers challenging enough, leaving substantial room for improvement. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated RewardsZhiming Lin, Kai Zhao, Sophie Zhang, Peilai Yu 等AAAI 2026 · 被引用 11 次
- Enhancing Character-Level Understanding in LLMs through Token Internal Structure LearningZhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu 等ACL 2025 · 被引用 7 次
- Mixture of Small and Large Models for Chinese Spelling CheckZiheng Qiao, Houquan Zhou, Zhenghua LiACL 2025 · 被引用 4 次
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li 等EMNLP 2024 · 被引用 2 次
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng 等EMNLP 2025
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 被引用 204 次
- Exploring and Adapting Chinese GPT to Pinyin Input MethodMinghuan Tan, Yong Dai, Duyu Tang, Zhangyin Feng 等ACL 2022 · 被引用 13 次
相关 Paper
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng 等EMNLP 2024 · 被引用 9 次
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling CheckXingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang 等ACL 2020 · 被引用 139 次
- Rethinking Masked Language Modeling for Chinese Spelling CorrectionHongqiu Wu, Shaohua Zhang, Yuchen Zhang, Hai ZhaoACL 2023 · 被引用 21 次
- ARM: An Alignment-and-Replacement Module for Chinese Spelling Check Based on LLMsChangchun Liu, Kai Zhang, Junzhe Jiang, Zirui Liu 等EMNLP 2024 · 被引用 3 次
- Chinese Spelling Correction as Rephrasing Language ModelLinfeng Liu, Hongqiu Wu, Hai ZhaoAAAI 2024 · 被引用 36 次
