Generalizing Clinical De-identification Models by Privacy-safe Data Augmentation using GPT-4
Woojin Kim, Sungeun Hahm, Jaejin Lee
摘要
De-identification (de-ID) refers to removing the association between a set of identifying data and the data subject. In clinical data management, the de-ID of Protected Health Information (PHI) is critical for patient confidentiality. However, state-of-the-art de-ID models show poor generalization on a new dataset. This is due to the difficulty of retaining training corpora. Additionally, labeling standards and the formats of patient records vary across different institutions. Our study addresses these issues by exploiting GPT-4 for data augmentation through one-shot and zero-shot prompts. Our approach effectively circumvents the problem of PHI leakage, ensuring privacy by redacting PHI before processing. To evaluate the effectiveness of our proposal, we conduct cross-dataset testing. The experimental result demonstrates significant improvements across three types of F1 scores.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information LossKiana Aghakasiri, Noopur Zambare, JoAnn Thai, Carrie Ye 等EMNLP 2025 · 被引用 1 次
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- MINIM: Privacy-Aware Minimal View for Agents via Trusted Local SanitizationHexuan Yu, Chaoyu Zhang, Heng Jin, Shanghao Shi 等ICML 2026
它引用的顶会 Paper2
相关 Paper
- Towards Injecting Medical Visual Knowledge into Multimodal LLMs at ScaleJunying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao 等EMNLP 2024 · 被引用 43 次
- Evaluating LLM-based Personal Information Extraction and CountermeasuresYupei Liu, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2025
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim 等EMNLP 2022 · 被引用 285 次
- Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT ModelsJunjie Chu, Zeyang Sha, Michael Backes, Yang ZhangEMNLP 2024 · 被引用 3 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
