Generalizing Clinical De-identification Models by Privacy-safe Data Augmentation using GPT-4
Woojin Kim, Sungeun Hahm, Jaejin Lee
Abstract
De-identification (de-ID) refers to removing the association between a set of identifying data and the data subject. In clinical data management, the de-ID of Protected Health Information (PHI) is critical for patient confidentiality. However, state-of-the-art de-ID models show poor generalization on a new dataset. This is due to the difficulty of retaining training corpora. Additionally, labeling standards and the formats of patient records vary across different institutions. Our study addresses these issues by exploiting GPT-4 for data augmentation through one-shot and zero-shot prompts. Our approach effectively circumvents the problem of PHI leakage, ensuring privacy by redacting PHI before processing. To evaluate the effectiveness of our proposal, we conduct cross-dataset testing. The experimental result demonstrates significant improvements across three types of F1 scores.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af3ab517-b74c-4cac-90b3-242972543bfbCited by top-tier papers3
- Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information LossKiana Aghakasiri, Noopur Zambare, JoAnn Thai, Carrie Ye et al.EMNLP 2025 · 1 citation
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- MINIM: Privacy-Aware Minimal View for Agents via Trusted Local SanitizationHexuan Yu, Chaoyu Zhang, Heng Jin, Shanghao Shi et al.ICML 2026
Builds on2
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai et al.EMNLP 2020 · 132 citations
- Clinical Reading Comprehension: A Thorough Analysis of the emrQA DatasetXiang Yue, Bernal Jimenez Gutierrez, Huan SunACL 2020 · 35 citations
Related papers
- Towards Injecting Medical Visual Knowledge into Multimodal LLMs at ScaleJunying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao et al.EMNLP 2024 · 43 citations
- Evaluating LLM-based Personal Information Extraction and CountermeasuresYupei Liu, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2025
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
- Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT ModelsJunjie Chu, Zeyang Sha, Michael Backes, Yang ZhangEMNLP 2024 · 3 citations
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
