Tail-to-Tail Non-Autoregressive Sequence Prediction for Chinese Grammatical Error Correction
Piji Li, Shuming Shi
摘要
We investigate the problem of Chinese Grammatical Error Correction (CGEC) and present a new framework named Tail-to-Tail (TtT) non-autoregressive sequence prediction to address the deep issues hidden in CGEC. Considering that most tokens are correct and can be conveyed directly from source to target, and the error positions can be estimated and corrected based on the bidirectional context information, thus we employ a BERTinitialized Transformer Encoder as the backbone model to conduct information modeling and conveying. Considering that only relying on the same position substitution cannot handle the variable-length correction cases, various operations such substitution, deletion, insertion, and local paraphrasing are required jointly. Therefore, a Conditional Random Fields (CRF) layer is stacked on the up tail to conduct non-autoregressive sequence prediction by modeling the token dependencies. Since most tokens are correct and easily to be predicted/conveyed to the target, then the models may suffer from a severe class imbalance issue. To alleviate this problem, focal loss penalty strategies are integrated into the loss functions. Moreover, besides the typical fix-length error correction datasets, we also construct a variable-length corpus to conduct experiments. Experimental results on standard datasets, especially on the variable-length datasets, demonstrate the effectiveness of TtT in terms of sentence-level Accuracy, Precision, Recall, and F1-Measure on tasks of error Detection and Correction 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Rethinking Masked Language Modeling for Chinese Spelling CorrectionHongqiu Wu, Shaohua Zhang, Yuchen Zhang, Hai ZhaoACL 2023 · 被引用 21 次
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng 等EMNLP 2024 · 被引用 9 次
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li 等EMNLP 2025
- Toward Robust Evaluation for Multilingual Grammatical Error Correction: Can Large Language Models Replace Human References?Alla Rozovskaya, Dan RothACL 2026
- CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion TypesZishan Guo, Linhao Yu, Minghui Xu, Renren Jin 等EMNLP 2023
它引用的顶会 Paper3
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 被引用 204 次
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling CheckXingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang 等ACL 2020 · 被引用 139 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
相关 Paper
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong 等ACL 2023 · 被引用 21 次
- MaskGEC: Improving Neural Grammatical Error Correction via Dynamic MaskingZewei Zhao, Houfeng WangAAAI 2020 · 被引用 71 次
- Sequence-to-Action: Grammatical Error Correction with Action Guided Sequence GenerationJiquan Li, Junliang Guo, Yongxin Zhu, Xin Sheng 等AAAI 2022 · 被引用 29 次
- ScholarGEC: Enhancing Controllability of Large Language Model for Chinese Academic Grammatical Error CorrectionZixiao Kong, Xianquan Wang, Shuanghong Shen, Keyu Zhu 等AAAI 2025 · 被引用 2 次
- Detection-Correction Structure via General Language Model for Grammatical Error CorrectionWei Li, Houfeng WangACL 2024 · 被引用 9 次
