MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model Score
Sunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy, Brian Corner, Hong Yu
摘要
This paper proposes a new natural language processing (NLP) application for identifying medical jargon terms potentially difficult for patients to comprehend from electronic health record (EHR) notes. We first present a novel and publicly available dataset with expertannotated medical jargon terms from 18K+ EHR note sentences (M edJ). Then, we introduce a novel medical jargon extraction (M edJEx) model which has been shown to outperform existing state-of-the-art NLP models. First, MedJEx improved the overall performance when it was trained on an auxiliary Wikipedia hyperlink span dataset, where hyperlink spans provide additional Wikipedia articles to explain the spans (or terms), and then fine-tuned on the annotated MedJ data. Secondly, we found that a contextualized masked language model score was beneficial for detecting domain-specific unfamiliar jargon terms. Moreover, our results show that training on the auxiliary Wikipedia hyperlink span datasets improved six out of eight biomedical named entity recognition benchmark datasets. MedJEx is publicly available 1 . UMLS QuickUMLS Weighted Score Feature Binary Feature Wiki_trained LM Tokenizer CRF Layer MLP MLP MLP Binary Feature Extraction Biomedical Concepts Term Weighting Initialize with trained weights WikiHyperlink Training Auxiliary Feature Extraction Target Model Input Hidden Weighted emission Emission Final emission WordFreq … exacerbated by his shock liver … … 'exacerbated by' 'shock' 'liver'
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Improving Summarization with Human EditsZonghai Yao, Benjamin J. Schloss, Sai P. SelvarajEMNLP 2023 · 被引用 6 次
- Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss InformationSunjae Kwon, Rishabh Garodia, Minhwa Lee, Zhichao Yang 等ACL 2023 · 被引用 3 次
- MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical DomainChao Jiang, Wei XuEMNLP 2024 · 被引用 3 次
- LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI HospitalsZonghai Yao, Hong YuACL 2026
它引用的顶会 Paper1
相关 Paper
- Understanding Jargon: Combining Extraction and Generation for Definition ModelingJie Huang, Hanyin Shao, Kevin Chen-Chuan Chang, Jinjun Xiong 等EMNLP 2022 · 被引用 11 次
- Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning RepresentationZixuan Zhang, Nikolaus Nova Parulian, Heng Ji, Ahmed Elsayed 等ACL 2021
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 被引用 56 次
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier 等ACL 2023 · 被引用 19 次
- Hierarchical Pretraining on Multimodal Electronic Health RecordsXiaochen Wang, Junyu Luo, Jiaqi Wang, Ziyi Yin 等EMNLP 2023 · 被引用 7 次
