SMedBERT: A Knowledge-Enhanced Pre-trained Language Model with Structured Semantics for Medical Text Mining
Taolin Zhang, Zerui Cai, Chengyu Wang, Minghui Qiu, Bite Yang, Xiaofeng He
摘要
Recently, the performance of Pre-trained Language Models (PLMs) has been significantly improved by injecting knowledge facts to enhance their abilities of language understanding. For medical domains, the background knowledge sources are especially useful, due to the massive medical terms and their complicated relations are difficult to understand in text. In this work, we introduce SMedBERT, a medical PLM trained on large-scale medical corpora, incorporating deep structured semantics knowledge from neighbours of linked-entity. In SMedBERT, the mention-neighbour hybrid attention is proposed to learn heterogeneousentity information, which infuses the semantic representations of entity types into the homogeneous neighbouring entity structure. Apart from knowledge integration as external features, we propose to employ the neighbors of linked-entities in the knowledge graph as additional global contexts of text mentions, allowing them to communicate via shared neighbors, thus enrich their semantic representations. Experiments demonstrate that SMedBERT significantly outperforms strong baselines in various knowledge-intensive Chinese medical tasks. It also improves the performance of other tasks such as question answering, question matching and natural language inference. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LasUIE: Unifying Information Extraction with Latent Adaptive Structure-aware Generative Language ModelHao Fei, Shengqiong Wu, Jingye Li, Bobo Li 等NeurIPS 2022 · 被引用 114 次
- UPPAM: A Unified Pre-training Architecture for Political Actor Modeling based on LanguageXinyi Mou, Zhongyu Wei, Qi Zhang, Xuanjing HuangACL 2023 · 被引用 6 次
- Learning Knowledge-Enhanced Contextual Language Representations for Domain Natural Language UnderstandingTaolin Zhang, Ruyao Xu, Chengyu Wang, Zhongjie Duan 等EMNLP 2023 · 被引用 1 次
它引用的顶会 Paper4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 被引用 294 次
- Infusing Disease Knowledge into BERT for Health Question Answering, Medical Inference and Disease Name RecognitionYun He, Ziwei Zhu, Yin Zhang, Qin Chen 等EMNLP 2020 · 被引用 103 次
- RikiNet: Reading Wikipedia Pages for Natural Question AnsweringDayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan 等ACL 2020 · 被引用 55 次
相关 Paper
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 被引用 56 次
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model AlignmentAndrey Sakhovskiy, Elena TutubalinaSIGIR 2025 · 被引用 2 次
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier 等ACL 2023 · 被引用 19 次
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 被引用 463 次
- Learning Conceptual-Contextual Embeddings for Medical TextXiao Zhang, Dejing Dou, Ji WuAAAI 2020 · 被引用 17 次
