Trustworthy Machine Learning through Data-Specific Indistinguishability
Hanshen Xiao, Zhen Yang, G. Edward Suh
摘要
This paper studies a range of AI/ML trust concepts, including memorization, data poisoning, and copyright, which can be modeled as constraints on the influence of data on a (trained) model, characterized by the outcome difference from a processing function (training algorithm). In this realm, we show that provable trust guarantees can be efficiently provided through a new framework termed Data-Specific Indistinguishability (DSI) to select trust-preserving randomization tightly aligning with targeted outcome differences, as a relaxation of the classic Input-Independent Indistinguishability (III). We establish both the theoretical and algorithmic foundations of DSI with the optimal multivariate Gaussian mechanism. We further show its applications to develop trustworthy deep learning with black-box optimizers. The experimental results on memorization mitigation, backdoor defense, and copyright protection show both the efficiency and effectiveness of the DSI noise mechanism.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu 等S&P 2018 · 被引用 867 次
相关 Paper
- Leave-one-out Distinguishability in Machine LearningJiayuan Ye, Anastasia Borovykh, Soufiane Hayou, Reza ShokriICLR 2024 · 被引用 21 次
- DLBox: New Model Training Framework for Protecting Training DataJaewon Hur, Juheon Yi, Cheolwoo Myung, Sangyun Kim 等NDSS 2025
- Safe and Robust Watermark Injection with a Single OoD ImageShuyang Yu, Junyuan Hong, Haobo Zhang, Haotao Wang 等ICLR 2024 · 被引用 4 次
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 被引用 543 次
- Bounding Training Data Reconstruction in Private (Deep) LearningChuan Guo, Brian Karrer, Kamalika Chaudhuri, Laurens van der MaatenICML 2022 · 被引用 66 次
