Please Tell Me More: Privacy Impact of Explainability through the Lens of Membership Inference Attack
Han Liu, Yuhao Wu, Zhiyuan Yu, Ning Zhang
Abstract
Explainability is increasingly recognized as an enabling technology for the broader adoption of machine learning (ML), particularly for safety-critical applications. This has given rise to explainable ML, which seeks to enhance the explainability of neural networks through the use of explanators. Yet, the pursuit for better explainability inadvertently leads to increased security and privacy risks. While there has been considerable research into the security risks of explainable ML, its potential privacy risks remain under-explored.To bridge this gap, we present a systematic study of privacy risks in explainable ML through the lens of membership inference. Building on the observation that, besides the accuracy of the model, robustness also exhibits observable differences among member samples and non-member samples, we develop a new membership inference attack. This attack extracts additional membership features from changes in model confidence under different levels of perturbations guided by the importance highlighted by the attribution maps in the explanators. Intuitively, perturbing important features generally results in a bigger loss in confidence for member samples. Using the member-non-member differences in both model performance and robustness, an attack model is trained to distinguish the membership. We evaluated our approach with seven popular explanators across various benchmark models and datasets. Our attack demonstrates there is non-trivial privacy leakage in current explainable ML methods. Furthermore, such leakage issue persists even if the attacker lacks the knowledge of training datasets or target model architectures. Lastly, we also found existing model and output-based defense mechanisms are not effective in mitigating this new attack.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c8dd191b-8903-49a7-a0c4-9acaf3e113afCited by top-tier papers10
- SoK: Unintended Interactions among Machine Learning Defenses and RisksVasisht Duddu, Sebastian Szyller, N. AsokanS&P 2024 · 6 citations
- In-Context Probing for Membership Inference in Fine-Tuned Language ModelsZhexi Lu, Hongliang Chi, Nathalie Baracaldo, Swanand Ravindra Kadhe et al.NDSS 2026 · 3 citations
- EcoLoRA: Communication-Efficient Federated Fine-Tuning of Large Language ModelsHan Liu, Ruoyao Wen, Srijith Nair, Jia Liu et al.EMNLP 2025 · 2 citations
- A Unified Defense Framework Against Membership Inference in Federated Learning via Distillation and Contribution-Aware AggregationLiwei Zhang, Linghui Li, Xiaotian Si, Ziduo Guo et al.NDSS 2026 · 1 citation
- CompLeak: Deep Learning Model Compression Exacerbates Privacy LeakageNa Li, Yansong Gao, Hongsheng Hu, Boyu Kuang et al.USENIX Security 2026
Related papers
- Privacy Risks of Securing Machine Learning Models against Adversarial ExamplesLiwei Song, Reza Shokri, Prateek MittalCCS 2019 · 293 citations
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 628 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
- Membership Inference Attacks and Defenses in Neural Network PruningXiaoyong Yuan, Lan ZhangUSENIX Security 2022
- Membership Leakage in Label-Only ExposuresZheng Li, Yang ZhangCCS 2021 · 185 citations
