Towards Real-World Writing Assistance: A Chinese Character Checking Benchmark with Faked and Misspelled Characters
Yinghui Li, Zishan Xu, Shaoshen Chen, Haojing Huang, Yangning Li, Shirong Ma, Yong Jiang, Zhongli Li, Qingyu Zhou, Hai-Tao Zheng, Ying Shen
摘要
Writing assistance aims to improve the correctness and quality of input texts, with character checking being crucial in detecting and correcting wrong characters. In the real world where handwriting occupies the vast majority, characters that humans get wrong include faked characters (i.e., untrue characters created due to writing errors) and misspelled characters (i.e., true characters used incorrectly due to spelling errors). However, existing datasets and related studies only focus on misspelled characters that can be represented by computer text encoding systems, thereby ignoring faked characters which are more common and difficult. To break through this dilemma, we present Visual-C 3 , a human-annotated Visual Chinese Character Checking dataset with faked and misspelled Chinese characters. To the best of our knowledge, Visual-C 3 is the first real-world visual and the largest humancrafted dataset for the Chinese character checking scenario. Additionally, we also propose and evaluate novel baseline methods on Visual-C 3 . Extensive empirical results and analyses show that Visual-C 3 is high-quality yet challenging. As the first study focusing on Chinese faked characters, the Visual-C 3 dataset and the baseline methods are publicly available at https: //github.com/THUKElab/Visual-C3 . * * indicates equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang 等AAAI 2025 · 被引用 138 次
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular ReasoningJiaru Zou, Soumya Roy, Vinay Kumar Verma, Ziyi Wang 等ICLR 2026 · 被引用 13 次
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang 等NeurIPS 2025 · 被引用 11 次
- From Diagrams to Code: Multilingual Programming with Visual DesignLinzheng Chai, Jian Yang, Shukai Liu, Wei Zhang 等ICML 2026
- CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error CorrectionJingheng Ye, Zishan Xu, Yinghui Li, Linlin Song 等ACL 2025
它引用的顶会 Paper3
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Rethinking Masked Language Modeling for Chinese Spelling CorrectionHongqiu Wu, Shaohua Zhang, Yuchen Zhang, Hai ZhaoACL 2023 · 被引用 21 次
- Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal TransportXuxin Cheng, Zhihong Zhu, Hongxiang Li, Yaowei Li 等AAAI 2024 · 被引用 19 次
相关 Paper
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 被引用 8 次
- A Training-free LLM-based Approach to General Chinese Character Error CorrectionHouquan Zhou, Bo Zhang, Zhenghua Li, Ming Yan 等ACL 2025
- CSCD-NS: a Chinese Spelling Check Dataset for Native SpeakersYong Hu, Fandong Meng, Jie ZhouACL 2024 · 被引用 9 次
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling CheckXingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang 等ACL 2020 · 被引用 139 次
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng 等EMNLP 2024 · 被引用 9 次
