Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence Modeling
Zhijun Wang, Xuebo Liu, Min Zhang
Abstract
Existing research generally treats Chinese character as a minimum unit for representation. However, such Chinese character representation will suffer two bottlenecks: 1) Learning bottleneck, the learning cannot benefit from its rich internal features (e.g., radicals and strokes); and 2) Parameter bottleneck, each individual character has to be represented by a unique vector. In this paper, we introduce a novel representation method for Chinese characters to break the bottlenecks, namely StrokeNet, which represents a Chinese character by a Latinized stroke sequence (e.g., "凹(concave)" to "ajaie" and "凸(convex)" to "aeaqe"). Specifically, StrokeNet maps each stroke to a specific Latin character, thus allowing similar Chinese characters to have similar Latin representations. With the introduction of StrokeNet to neural machine translation (NMT), many powerful but not applicable techniques to non-Latin languages (e.g., shared subword vocabulary learning and ciphertextbased data augmentation) can now be perfectly implemented. Experiments on the widelyused NIST Chinese-English, WMT17 Chinese-English and IWSLT17 Japanese-English NMT tasks show that StrokeNet can provide a significant performance boost over the strong baselines with fewer model parameters, achieving 26.5 BLEU on the WMT17 Chinese-English task which is better than any previously reported results without using monolingual data. Code and scripts are freely available at https: //github.com/zjwang21/StrokeNet . * Co-first and Corresponding Author Zh 布 什和沙龙举行了 会谈 En Bush held a talk with Sharon Zh (Stroke) etasa taea teatoaie oodatot etcto ootetneea ttaeer hr tneelo oyottoottn Table 1: StrokeNet represents a Chinese character by a Latinized stroke sequence. For example, "布" to "etasa" and "了" to "hr".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 902b7b98-8e41-4d22-90f4-971f67e6315dCited by top-tier papers3
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao et al.EMNLP 2022 · 20 citations
- kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine TranslationShudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li et al.ACL 2023 · 14 citations
- Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech TranslationTengfei Yu, Xuebo Liu, Liang Ding, Kehai Chen et al.ACL 2024
Builds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
- AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationYong Cheng, Lu Jiang, Wolfgang Macherey, Jacob EisensteinACL 2020 · 105 citations
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 97 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
Related papers
- Stroke Extraction of Chinese Character Based on Deep Structure Deformable Image RegistrationMeng Li, Yahan Yu, Yi Yang, Guanghao Ren et al.AAAI 2023 · 7 citations
- Simplify the Usage of Lexicon in Chinese NERRuotian Ma, Minlong Peng, Qi Zhang, Zhongyu Wei et al.ACL 2020 · 286 citations
- StrokeGAN: Reducing Mode Collapse in Chinese Font Generation via Stroke EncodingJinshan Zeng, Qi Chen, Yunxin Liu, Mingwen Wang et al.AAAI 2021 · 68 citations
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin InformationZijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng et al.ACL 2021
- 2kenize: Tying Subword Sequences for Chinese Script ConversionPranav A, Isabelle AugensteinACL 2020 · 4 citations
