Are Attribute Inference Attacks Just Imputation?
Bargav Jayaraman, David Evans
Abstract
Models can expose sensitive information about their training data. In an attribute inference attack, an adversary has partial knowledge of some training records and access to a model trained on those records, and infers the unknown values of a sensitive feature of those records. We study a fine-grained variant of attribute inference we call sensitive value inference, where the adversary's goal is to identify with high confidence some records from a candidate set where the unknown attribute has a particular sensitive value. We explicitly compare attribute inference with data imputation that captures the training distribution statistics, under various assumptions about the training data available to the adversary. Our main conclusions are: (1) previous attribute inference methods do not reveal more about the training data from the model than can be inferred by an adversary without access to the trained model, but with the same knowledge of the underlying distribution as needed to train the attribute inference attack; (2) black-box attribute inference attacks rarely learn anything that cannot be learned without the model; but (3) white-box attacks, which we introduce and evaluate in the paper, can reliably identify some records with the sensitive value attribute that would not be predicted without having access to the model. Furthermore, we show that proposed defenses such as differentially private training and removing vulnerable records from training do not mitigate this privacy risk. The code for our experiments is available at https://github.com/bargavj/EvaluatingDPML . Contributions. To better understand attribute inference risks, we consider threat models where an adversary has limited prior knowledge of the training distribution (Section 3.1) and study a finergrained notion of attribute inference that considers the privacy risk of identifying, with high confidence, individuals with a particular sensitive attribute values from a candidate set (Section 3.2). We propose a novel white-box attack that identifies neurons in a model that are most correlated with the sensitive value for a target attribute (Section 5). We perform extensive experimental evaluation sensitive value inference on two large real world data sets with both imputation and black-box attribute inference attacks (Section 7) and with our novel white-box attacks (Section 8). Key findings. Our experiments show that trained models leak considerable information about the underlying training distribution which can be exploited to infer sensitive attributes about individuals. While prior attribute inference attacks do not learn anything from the model that could not be learned without it, our white-box attacks are able to confidently infer sensitive value records, even 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 37 citations
- Do SSL Models Have Déjà Vu? A Case of Unintended Memorization in Self-supervised LearningCasey Meehan, Florian Bordes, Pascal Vincent, Kamalika Chaudhuri et al.NeurIPS 2023 · 26 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- SoK: Unintended Interactions among Machine Learning Defenses and RisksVasisht Duddu, Sebastian Szyller, N. AsokanS&P 2024 · 6 citations
- Analyzing Inference Privacy Risks Through Gradients In Machine LearningZhuohang Li, Andrew Lowy, Jing Liu, Toshiaki Koike-Akino et al.CCS 2024 · 5 citations
Builds on6
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Are Your Sensitive Attributes Private? Novel Model Inversion Attribute Inference Attacks on Classification ModelsShagufta Mehnaz, Sayanton V. Dibbo, Ehsanul Kabir, Ninghui Li et al.USENIX Security 2022
- ML-Doctor: Holistic Risk Assessment of Inference Attacks Against Machine Learning ModelsYugeng Liu, Rui Wen, Xinlei He, Ahmed Salem et al.USENIX Security 2022
Related papers
- Disparate Privacy Vulnerability: Targeted Attribute Inference Attacks and DefensesEhsanul Kabir, Lucas Craig, Shagufta MehnazUSENIX Security 2025
- Group Property Inference Attacks Against Graph Neural NetworksXiuling Wang, Wendy Hui WangCCS 2022 · 31 citations
- Can we estimate privacy vulnerability of individual records? Towards Mitigating Attribute Inference Attacks on ML ModelsEhsanul Kabir, Najrin Sultana, Ninghui Li, Shagufta MehnazUSENIX Security 2026
- Leakage of Dataset Properties in Multi-Party Machine LearningWanrong Zhang, Shruti Tople, Olga OhrimenkoUSENIX Security 2021 · 92 citations
- Property Inference from PoisoningSaeed Mahloujifar, Esha Ghosh, Melissa ChaseS&P 2022 · 96 citations
