Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss
Kiana Aghakasiri, Noopur Zambare, JoAnn Thai, Carrie Ye, Mayur Mehta, J. Ross Mitchell, Mohamed Abdalla
摘要
De-identification in the healthcare setting is an application of NLP where automated algorithms are used to remove personally identifying information of patients (and, sometimes, providers). With the recent rise of generative large language models (LLMs), there has been a corresponding rise in the number of papers that apply LLMs to de-identification. Although these approaches often report near-perfect results, significant challenges concerning reproducibility and utility of the research papers persist. This paper identifies three key limitations in the current literature: inconsistent reporting metrics hindering direct comparisons, the inadequacy of traditional classification metrics in capturing errors which LLMs may be more prone to (i.e., altering clinically relevant information), and lack of manual validation of automated metrics which aim to quantify these errors. To address these issues, we first present a survey of LLM-based de-identification research, highlighting the heterogeneity in reporting standards. Second, we evaluated a diverse set of models to quantify the extent of inappropriate removal of clinical information. Next, we conduct a manual validation of an existing evaluation metric to measure the removal of clinical information, employing clinical experts to assess their efficacy. We highlight poor performance and describe the inherent limitations of such metrics in identifying clinically significant changes. Lastly, we propose a novel methodology for the detection of clinically relevant information removal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsHye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. WallaceEMNLP 2023 · 被引用 11 次
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung 等ACL 2026 · 被引用 9 次
- Factuality of Large Language Models: A SurveyYuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu 等EMNLP 2024 · 被引用 23 次
- Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians' Evaluation of Large Language Models for Smoking Cessation InterventionsPaul Calle, Ruosi Shao, Yunlong Liu, Emily T. Hébert 等CHI 2024 · 被引用 20 次
- Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation GenerationAneta Zugecova, Dominik Macko, Ivan Srba, Róbert Móro 等ACL 2025 · 被引用 18 次
