ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
Zijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng, Xiang Ao, Qing He, Fei Wu, Jiwei Li
摘要
Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT, which incorporates both the glyph and pinyin information of Chinese characters into language model pretraining. The glyph embedding is obtained based on different fonts of a Chinese character, being able to capture character semantics from the visual features, and the pinyin embedding characterizes the pronunciation of Chinese characters, which handles the highly prevalent heteronym phenomenon in Chinese (the same character has different pronunciations with different meanings). Pretrained on large-scale unlabeled Chinese corpus, the proposed Chine-seBERT model yields significant performance boost over baseline models with fewer training steps. The proposed model achieves new SOTA performances on a wide range of Chinese NLP tasks,including machine reading comprehension, natural language inference, text classification, sentence pair matching, and competitive performances in named entity recognition and word segmentation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingHui Su, Weiwei Shi, Xiaoyu Shen, Xiao Zhou 等ACL 2022 · 被引用 38 次
- GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language ModelsJian Ma, Yonglin Deng, Chen Chen, Nanyang Du 等AAAI 2025 · 被引用 28 次
- Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and GranularityJiahao Li, Quan Wang, Zhendong Mao, Junbo Guo 等EMNLP 2022 · 被引用 19 次
- Exploring and Adapting Chinese GPT to Pinyin Input MethodMinghuan Tan, Yong Dai, Duyu Tang, Zhangyin Feng 等ACL 2022 · 被引用 13 次
- Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence ModelingZhijun Wang, Xuebo Liu, Min ZhangEMNLP 2022 · 被引用 10 次
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng 等AAAI 2020 · 被引用 885 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He 等ICLR 2020 · 被引用 391 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
相关 Paper
- PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling CheckLi Huang, Junjie Li, Weiwei Jiang, Zhiyu Zhang 等ACL 2021
- Disentangled Phonetic Representation for Chinese Spelling CorrectionZihong Liang, Xiaojun Quan, Qifan WangACL 2023 · 被引用 12 次
- Lexicon Enhanced Chinese Sequence Labeling Using BERT AdapterWei Liu, Xiyan Fu, Yue Zhang, Wenming XiaoACL 2021
- 2kenize: Tying Subword Sequences for Chinese Script ConversionPranav A, Isabelle AugensteinACL 2020 · 被引用 4 次
- Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics GraphYanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang 等ACL 2022 · 被引用 11 次
