MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model Score
Sunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy, Brian Corner, Hong Yu
Abstract
This paper proposes a new natural language processing (NLP) application for identifying medical jargon terms potentially difficult for patients to comprehend from electronic health record (EHR) notes. We first present a novel and publicly available dataset with expertannotated medical jargon terms from 18K+ EHR note sentences (M edJ). Then, we introduce a novel medical jargon extraction (M edJEx) model which has been shown to outperform existing state-of-the-art NLP models. First, MedJEx improved the overall performance when it was trained on an auxiliary Wikipedia hyperlink span dataset, where hyperlink spans provide additional Wikipedia articles to explain the spans (or terms), and then fine-tuned on the annotated MedJ data. Secondly, we found that a contextualized masked language model score was beneficial for detecting domain-specific unfamiliar jargon terms. Moreover, our results show that training on the auxiliary Wikipedia hyperlink span datasets improved six out of eight biomedical named entity recognition benchmark datasets. MedJEx is publicly available 1 . UMLS QuickUMLS Weighted Score Feature Binary Feature Wiki_trained LM Tokenizer CRF Layer MLP MLP MLP Binary Feature Extraction Biomedical Concepts Term Weighting Initialize with trained weights WikiHyperlink Training Auxiliary Feature Extraction Target Model Input Hidden Weighted emission Emission Final emission WordFreq … exacerbated by his shock liver … … 'exacerbated by' 'shock' 'liver'
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9de07745-0611-4bb9-ba07-69a5d2efb380Cited by top-tier papers4
- Improving Summarization with Human EditsZonghai Yao, Benjamin J. Schloss, Sai P. SelvarajEMNLP 2023 · 6 citations
- Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss InformationSunjae Kwon, Rishabh Garodia, Minhwa Lee, Zhichao Yang et al.ACL 2023 · 3 citations
- MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical DomainChao Jiang, Wei XuEMNLP 2024 · 3 citations
- LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI HospitalsZonghai Yao, Hong YuACL 2026
Builds on1
Related papers
- Understanding Jargon: Combining Extraction and Generation for Definition ModelingJie Huang, Hanyin Shao, Kevin Chen-Chuan Chang, Jinjun Xiong et al.EMNLP 2022 · 11 citations
- Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning RepresentationZixuan Zhang, Nikolaus Nova Parulian, Heng Ji, Ahmed Elsayed et al.ACL 2021
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 56 citations
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier et al.ACL 2023 · 19 citations
- Hierarchical Pretraining on Multimodal Electronic Health RecordsXiaochen Wang, Junyu Luo, Jiaqi Wang, Ziyi Yin et al.EMNLP 2023 · 7 citations
