2kenize: Tying Subword Sequences for Chinese Script Conversion
Pranav A, Isabelle Augenstein
Abstract
Simplified Chinese to Traditional Chinese character conversion is a common preprocessing step in Chinese NLP. Despite this, current approaches have poor performance because they do not take into account that a simplified Chinese character can correspond to multiple traditional characters. Here, we propose a model that can disambiguate between mappings and convert between the two scripts. The model is based on subword segmentation, two language models, as well as a method for mapping between subword sequences. We further construct benchmark datasets for topic classification and script conversion. Our proposed method outperforms previous Chinese Character conversion approaches by 6 points in accuracy. These results are further confirmed in a downstream application, where 2kenize is used to convert pretraining dataset for topic classification. An error analysis reveals that our method's particular strengths are in dealing with code mixing and named entities. The code and dataset is available at https://github .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0aa3d78-4bf2-4f1c-9a82-cfbb482f26d9Builds on2
Related papers
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin InformationZijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng et al.ACL 2021
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer et al.ACL 2024
- Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningHaiyang Yu, Xiaocong Wang, Bin Li, Xiangyang XueICCV 2023 · 43 citations
- BTS: A Bi-lingual Benchmark for Text Segmentation in the WildXixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang et al.CVPR 2022 · 14 citations
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng et al.EMNLP 2024 · 9 citations
