Trustworthy Machine Learning through Data-Specific Indistinguishability
Hanshen Xiao, Zhen Yang, G. Edward Suh
Abstract
This paper studies a range of AI/ML trust concepts, including memorization, data poisoning, and copyright, which can be modeled as constraints on the influence of data on a (trained) model, characterized by the outcome difference from a processing function (training algorithm). In this realm, we show that provable trust guarantees can be efficiently provided through a new framework termed Data-Specific Indistinguishability (DSI) to select trust-preserving randomization tightly aligning with targeted outcome differences, as a relaxation of the classic Input-Independent Indistinguishability (III). We establish both the theoretical and algorithmic foundations of DSI with the optimal multivariate Gaussian mechanism. We further show its applications to develop trustworthy deep learning with black-box optimizers. The experimental results on memorization mitigation, backdoor defense, and copyright protection show both the efficiency and effectiveness of the DSI noise mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 239d1ac8-100c-4814-9a22-3b6b7635b0f4Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
Related papers
- Leave-one-out Distinguishability in Machine LearningJiayuan Ye, Anastasia Borovykh, Soufiane Hayou, Reza ShokriICLR 2024 · 21 citations
- DLBox: New Model Training Framework for Protecting Training DataJaewon Hur, Juheon Yi, Cheolwoo Myung, Sangyun Kim et al.NDSS 2025
- Safe and Robust Watermark Injection with a Single OoD ImageShuyang Yu, Junyuan Hong, Haobo Zhang, Haotao Wang et al.ICLR 2024 · 4 citations
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 543 citations
- Bounding Training Data Reconstruction in Private (Deep) LearningChuan Guo, Brian Karrer, Kamalika Chaudhuri, Laurens van der MaatenICML 2022 · 66 citations
