Reducing information dependency does not cause training data privacy. Adversarially non-robust features do.
Rasmus Torp, Shailen Smith, Adam Breuer
Abstract
In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instead caused by a tunable connection to adversarial robustness. We begin by presenting three surprising results: (1) recent defenses that inhibit reconstruction by Model Inversion Attacks (MIAs), which evaluate leakage under an idealized attacker, do not reduce standard measures of information dependency (HSIC); (2) models that maximally memorize their training datasets remain robust to MIA reconstruction; and (3) models trained without seeing 97% of the training pixels, where recent information-theoretic bounds give arbitrarily strong privacy guarantees under standard assumptions, can still be devastatingly reconstructed by MIA. To explain these findings, we provide causal evidence that privacy under MIA arises from what the adversarial examples literature calls "non-robust" features (generalizable but imperceptible and unstable features). We further show that recent MIA defenses obtain their privacy improvements by unintentionally shifting models toward such features. To establish this causal relationship, we introduce Anti Adversarial Training (AT-AT ), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses. Our results revise the prevailing understanding of training data exposure and reveal a new privacy-robustness tradeoff. * Equal contribution. † Breuer Lab gratefully acknowledges the support of the OpenAI Cybersecurity Grant. Replication code is available at https://github.com/BreuerLabs/Anti-Adversarial-Training Published as a conference paper at ICLR 2026 tion domain. Unlike more conservative attacks that merely probe for the presence of some exposed training examples, MIAs test the degree to which a powerful attacker armed with white-box model access and significant computational and data resources can reconstruct information from the training examples. While early MIAs applied SGD to directly optimize reconstructions of individual training examples (Fredrikson et al., 2015) , contemporary MIAs use sophisticated gradient techniques and external data to attempt to infer the full set of class-level characteristics for each class in the target model's training data (Qiu et al., 2024; Struppek et al., 2022; Haim et al., 2022) . However, while MIAs lower-bound the extent to which contemporary vision models expose their training data to reconstruction, they do not explain what aspect of a model encodes this vulnerability, or how to prevent it. This raises a fundamental question: What properties of learned representations encode vulnerability to training data leakage and reconstruction, and how can they be controlled? ⋄ ⋄ ⋄ Understanding these properties has far-reaching implications for learning that extend beyond the privacy domain. For example, influential recent work conjectures that obtaining stronger learning performance may require more extensive model-to-training-data dependencies, including rote memorization and exposure of the training set, not only for vanilla CNNs (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on41
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
Related papers
- No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural NetworksYehonathan Refael, Guy Smorodinsky, Ofir Lindenbaum, Itay SafranICLR 2026 · 1 citation
- Model Inversion Robustness: Can Transfer Learning Help?Sy-Tuyen Ho, Koh Jun Hao, Keshigeyan Chandrasegaran, Ngoc-Bao Nguyen et al.CVPR 2024
- Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature FilteringHongyao Yu, Yixiang Qiu, Hao Fang, Tianqu Zhuang et al.KDD 2026 · 2 citations
- Bilateral Dependency Optimization: Defending Against Model-inversion AttacksXiong Peng, Feng Liu, Jingfeng Zhang, Long Lan et al.KDD 2022 · 20 citations
- Membership Inference Attacks are Easier on Difficult ProblemsAvital Shafran, Shmuel Peleg, Yedid HoshenICCV 2021 · 24 citations
