ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
Zijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng, Xiang Ao, Qing He, Fei Wu, Jiwei Li
Abstract
Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT, which incorporates both the glyph and pinyin information of Chinese characters into language model pretraining. The glyph embedding is obtained based on different fonts of a Chinese character, being able to capture character semantics from the visual features, and the pinyin embedding characterizes the pronunciation of Chinese characters, which handles the highly prevalent heteronym phenomenon in Chinese (the same character has different pronunciations with different meanings). Pretrained on large-scale unlabeled Chinese corpus, the proposed Chine-seBERT model yields significant performance boost over baseline models with fewer training steps. The proposed model achieves new SOTA performances on a wide range of Chinese NLP tasks,including machine reading comprehension, natural language inference, text classification, sentence pair matching, and competitive performances in named entity recognition and word segmentation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df563a57-6803-4801-b3cf-52e984327968Cited by top-tier papers9
- RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingHui Su, Weiwei Shi, Xiaoyu Shen, Xiao Zhou et al.ACL 2022 · 38 citations
- GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language ModelsJian Ma, Yonglin Deng, Chen Chen, Nanyang Du et al.AAAI 2025 · 28 citations
- Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and GranularityJiahao Li, Quan Wang, Zhendong Mao, Junbo Guo et al.EMNLP 2022 · 19 citations
- Exploring and Adapting Chinese GPT to Pinyin Input MethodMinghuan Tan, Yong Dai, Duyu Tang, Zhangyin Feng et al.ACL 2022 · 13 citations
- Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence ModelingZhijun Wang, Xuebo Liu, Min ZhangEMNLP 2022 · 10 citations
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
Related papers
- PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling CheckLi Huang, Junjie Li, Weiwei Jiang, Zhiyu Zhang et al.ACL 2021
- Disentangled Phonetic Representation for Chinese Spelling CorrectionZihong Liang, Xiaojun Quan, Qifan WangACL 2023 · 12 citations
- Lexicon Enhanced Chinese Sequence Labeling Using BERT AdapterWei Liu, Xiyan Fu, Yue Zhang, Wenming XiaoACL 2021
- 2kenize: Tying Subword Sequences for Chinese Script ConversionPranav A, Isabelle AugensteinACL 2020 · 4 citations
- Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics GraphYanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang et al.ACL 2022 · 11 citations
