2kenize: Tying Subword Sequences for Chinese Script Conversion
Pranav A, Isabelle Augenstein
摘要
Simplified Chinese to Traditional Chinese character conversion is a common preprocessing step in Chinese NLP. Despite this, current approaches have poor performance because they do not take into account that a simplified Chinese character can correspond to multiple traditional characters. Here, we propose a model that can disambiguate between mappings and convert between the two scripts. The model is based on subword segmentation, two language models, as well as a method for mapping between subword sequences. We further construct benchmark datasets for topic classification and script conversion. Our proposed method outperforms previous Chinese Character conversion approaches by 6 points in accuracy. These results are further confirmed in a downstream application, where 2kenize is used to convert pretraining dataset for topic classification. An error analysis reveals that our method's particular strengths are in dealing with code mixing and named entities. The code and dataset is available at https://github .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin InformationZijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng 等ACL 2021
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer 等ACL 2024
- Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningHaiyang Yu, Xiaocong Wang, Bin Li, Xiangyang XueICCV 2023 · 被引用 43 次
- BTS: A Bi-lingual Benchmark for Text Segmentation in the WildXixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang 等CVPR 2022 · 被引用 14 次
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng 等EMNLP 2024 · 被引用 9 次
