An Investigation of Memorization Risk in Healthcare Foundation Models
Sana Tonekaboni, Lena Stempfle, Adibvafa Fallahpour, Walter Gerych, Marzyeh Ghassemi
摘要
Foundation models trained on large-scale de-identified electronic health records (EHRs) hold promise for clinical applications. However, their capacity to memorize patient information raises important privacy concerns. In this work, we introduce a suite of black-box evaluation tests to assess privacy-related memorization risks in foundation models trained on structured EHR data. Our framework includes methods for probing memorization at both the embedding and generative levels, and aims to distinguish between model generalization and harmful memorization in clinically relevant settings. We contextualize memorization in terms of its potential to compromise patient privacy, particularly for vulnerable subgroups. We validate our approach on a publicly available EHR foundation model and release an opensource toolkit to facilitate reproducible and collaborative privacy assessments in healthcare AI. 2 * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 被引用 1,736 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
相关 Paper
- MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMsZhan Qu, Michael FärberACL 2026 · 被引用 1 次
- Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion ModelsQiao Li, Xiaomeng Fu, Xi Wang, Jin Liu 等ACM MM 2024 · 被引用 6 次
- Unveiling Memorization in Code ModelsZhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi 等ICSE 2024 · 被引用 33 次
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory ProbingJinhua Yin, Peiru Yang, Chen Yang, Huili Wang 等NeurIPS 2025 · 被引用 4 次
- Semantics Versus Identity: A Divide-and-Conquer Approach Towards Adjustable Medical Image De-IdentificationYuan Tian, Shuo Wang, Rongzhao Zhang, Zijian Chen 等ICCV 2025 · 被引用 3 次
