Please Tell Me More: Privacy Impact of Explainability through the Lens of Membership Inference Attack
Han Liu, Yuhao Wu, Zhiyuan Yu, Ning Zhang
摘要
Explainability is increasingly recognized as an enabling technology for the broader adoption of machine learning (ML), particularly for safety-critical applications. This has given rise to explainable ML, which seeks to enhance the explainability of neural networks through the use of explanators. Yet, the pursuit for better explainability inadvertently leads to increased security and privacy risks. While there has been considerable research into the security risks of explainable ML, its potential privacy risks remain under-explored.To bridge this gap, we present a systematic study of privacy risks in explainable ML through the lens of membership inference. Building on the observation that, besides the accuracy of the model, robustness also exhibits observable differences among member samples and non-member samples, we develop a new membership inference attack. This attack extracts additional membership features from changes in model confidence under different levels of perturbations guided by the importance highlighted by the attribution maps in the explanators. Intuitively, perturbing important features generally results in a bigger loss in confidence for member samples. Using the member-non-member differences in both model performance and robustness, an attack model is trained to distinguish the membership. We evaluated our approach with seven popular explanators across various benchmark models and datasets. Our attack demonstrates there is non-trivial privacy leakage in current explainable ML methods. Furthermore, such leakage issue persists even if the attacker lacks the knowledge of training datasets or target model architectures. Lastly, we also found existing model and output-based defense mechanisms are not effective in mitigating this new attack.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper10
- SoK: Unintended Interactions among Machine Learning Defenses and RisksVasisht Duddu, Sebastian Szyller, N. AsokanS&P 2024 · 被引用 6 次
- In-Context Probing for Membership Inference in Fine-Tuned Language ModelsZhexi Lu, Hongliang Chi, Nathalie Baracaldo, Swanand Ravindra Kadhe 等NDSS 2026 · 被引用 3 次
- EcoLoRA: Communication-Efficient Federated Fine-Tuning of Large Language ModelsHan Liu, Ruoyao Wen, Srijith Nair, Jia Liu 等EMNLP 2025 · 被引用 2 次
- A Unified Defense Framework Against Membership Inference in Federated Learning via Distillation and Contribution-Aware AggregationLiwei Zhang, Linghui Li, Xiaotian Si, Ziduo Guo 等NDSS 2026 · 被引用 1 次
- CompLeak: Deep Learning Model Compression Exacerbates Privacy LeakageNa Li, Yansong Gao, Hongsheng Hu, Boyu Kuang 等USENIX Security 2026
相关 Paper
- Privacy Risks of Securing Machine Learning Models against Adversarial ExamplesLiwei Song, Reza Shokri, Prateek MittalCCS 2019 · 被引用 293 次
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 被引用 628 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
- Membership Inference Attacks and Defenses in Neural Network PruningXiaoyong Yuan, Lan ZhangUSENIX Security 2022
- Membership Leakage in Label-Only ExposuresZheng Li, Yang ZhangCCS 2021 · 被引用 185 次
