Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling
Peijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie, Meishan Zhang, Min Zhang
摘要
Boundary information is critical for various Chinese language processing tasks, such as word segmentation, part-of-speech tagging, and named entity recognition. Previous studies usually resorted to the use of a high-quality external lexicon, where lexicon items can offer explicit boundary information. However, to ensure the quality of the lexicon, great human effort is always necessary, which has been generally ignored. In this work, we suggest unsupervised statistical boundary information instead, and propose an architecture to encode the information directly into pre-trained language models, resulting in Boundary-Aware BERT (BABERT). We apply BABERT for feature induction of Chinese sequence labeling tasks. Experimental results on ten benchmarks of Chinese sequence labeling demonstrate that BABERT can provide consistent improvements on all datasets. In addition, our method can complement previous supervised lexicon exploration, where further improvements can be achieved when integrated with external lexicon information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Improving Chinese Word Segmentation with Wordhood Memory NetworksYuanhe Tian, Yan Song, Fei Xia, Tong Zhang 等ACL 2020 · 被引用 95 次
- Entity Enhanced BERT Pre-training for Chinese NERChen Jia, Yuefeng Shi, Qinrong Yang, Yue ZhangEMNLP 2020 · 被引用 59 次
- Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-way Attentions of Auto-analyzed KnowledgeYuanhe Tian, Yan Song, Xiang Ao, Fei Xia 等ACL 2020 · 被引用 50 次
相关 Paper
- Lexicon Enhanced Chinese Sequence Labeling Using BERT AdapterWei Liu, Xiyan Fu, Yue Zhang, Wenming XiaoACL 2021
- SSMI: Semantic Similarity and Mutual Information Maximization Based Enhancement for Chinese NERPengnian Qi, Biao QinAAAI 2023 · 被引用 10 次
- SENCR: A Span Enhanced Two-Stage Network with Counterfactual Rethinking for Chinese NERHang Zheng, Qingsong Li, Shen Chen, Yuxuan Liang 等AAAI 2024 · 被引用 8 次
- Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics GraphYanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang 等ACL 2022 · 被引用 11 次
- Simplify the Usage of Lexicon in Chinese NERRuotian Ma, Minlong Peng, Qi Zhang, Zhongyu Wei 等ACL 2020 · 被引用 286 次
