Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence Modeling
Zhijun Wang, Xuebo Liu, Min Zhang
摘要
Existing research generally treats Chinese character as a minimum unit for representation. However, such Chinese character representation will suffer two bottlenecks: 1) Learning bottleneck, the learning cannot benefit from its rich internal features (e.g., radicals and strokes); and 2) Parameter bottleneck, each individual character has to be represented by a unique vector. In this paper, we introduce a novel representation method for Chinese characters to break the bottlenecks, namely StrokeNet, which represents a Chinese character by a Latinized stroke sequence (e.g., "凹(concave)" to "ajaie" and "凸(convex)" to "aeaqe"). Specifically, StrokeNet maps each stroke to a specific Latin character, thus allowing similar Chinese characters to have similar Latin representations. With the introduction of StrokeNet to neural machine translation (NMT), many powerful but not applicable techniques to non-Latin languages (e.g., shared subword vocabulary learning and ciphertextbased data augmentation) can now be perfectly implemented. Experiments on the widelyused NIST Chinese-English, WMT17 Chinese-English and IWSLT17 Japanese-English NMT tasks show that StrokeNet can provide a significant performance boost over the strong baselines with fewer model parameters, achieving 26.5 BLEU on the WMT17 Chinese-English task which is better than any previously reported results without using monolingual data. Code and scripts are freely available at https: //github.com/zjwang21/StrokeNet . * Co-first and Corresponding Author Zh 布 什和沙龙举行了 会谈 En Bush held a talk with Sharon Zh (Stroke) etasa taea teatoaie oodatot etcto ootetneea ttaeer hr tneelo oyottoottn Table 1: StrokeNet represents a Chinese character by a Latinized stroke sequence. For example, "布" to "etasa" and "了" to "hr".
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao 等EMNLP 2022 · 被引用 20 次
- kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine TranslationShudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li 等ACL 2023 · 被引用 14 次
- Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech TranslationTengfei Yu, Xuebo Liu, Liang Ding, Kehai Chen 等ACL 2024
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta 等ICLR 2022 · 被引用 198 次
- AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationYong Cheng, Lu Jiang, Wolfgang Macherey, Jacob EisensteinACL 2020 · 被引用 105 次
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 被引用 97 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
相关 Paper
- Stroke Extraction of Chinese Character Based on Deep Structure Deformable Image RegistrationMeng Li, Yahan Yu, Yi Yang, Guanghao Ren 等AAAI 2023 · 被引用 7 次
- Simplify the Usage of Lexicon in Chinese NERRuotian Ma, Minlong Peng, Qi Zhang, Zhongyu Wei 等ACL 2020 · 被引用 286 次
- StrokeGAN: Reducing Mode Collapse in Chinese Font Generation via Stroke EncodingJinshan Zeng, Qi Chen, Yunxin Liu, Mingwen Wang 等AAAI 2021 · 被引用 68 次
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin InformationZijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng 等ACL 2021
- 2kenize: Tying Subword Sequences for Chinese Script ConversionPranav A, Isabelle AugensteinACL 2020 · 被引用 4 次
