SoK: Unintended Interactions among Machine Learning Defenses and Risks
Vasisht Duddu, Sebastian Szyller, N. Asokan
Abstract
Machine learning (ML) models cannot neglect risks to security, privacy, and fairness. Several defenses have been proposed to mitigate such risks. When a defense is effective in mitigating one risk, it may correspond to increased or decreased susceptibility to other risks. Existing research lacks an effective framework to recognize and explain these unintended interactions. We present such a framework, based on the conjecture that overfitting and memorization underlie unintended interactions. We survey existing literature on unintended interactions, accommodating them within our framework. We use our framework to conjecture two previously unexplored interactions, and empirically validate them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ''People can change, and patterns can be broken'': Contextualizing Tradeoffs in Automated Decision-Making SystemsRabeya Bosri, Anna Harbluk Lorimer, Afrida Hossain, Vasisht Duddu et al.CCS 2026
- SoK: Colluding Adversaries in Machine Learning PipelinesVasisht Duddu, Lipeng He, Asim Waheed, N. AsokanUSENIX Security 2026
Builds on87
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 1,822 citations
Related papers
- The Privacy Onion Effect: Memorization is RelativeNicholas Carlini, Matthew Jagielski, Chiyuan Zhang, Nicolas Papernot et al.NeurIPS 2022 · 175 citations
- Overlearning Reveals Sensitive AttributesCongzheng Song, Vitaly ShmatikovICLR 2020 · 177 citations
- Can we estimate privacy vulnerability of individual records? Towards Mitigating Attribute Inference Attacks on ML ModelsEhsanul Kabir, Najrin Sultana, Ninghui Li, Shagufta MehnazUSENIX Security 2026
- Measuring Forgetting of Memorized Training ExamplesMatthew Jagielski, Om Thakkar, Florian Tramèr, Daphne Ippolito et al.ICLR 2023 · 15 citations
- Learning Robust and Privacy-Preserving Representations via Information TheoryBinghui Zhang, Sayedeh Leila Noorbakhsh, Yun Dong, Yuan Hong et al.AAAI 2025 · 4 citations
