Label differential privacy and private training data release
Róbert Istvan Busa-Fekete, Andrés Muñoz Medina, Umar Syed, Sergei Vassilvitskii
摘要
We study differentially private mechanisms for sharing training data in machine learning settings. Our goal is to enable learning of an accurate predictive model while protecting the privacy of each user's label. Previous work established privacy guarantees that assumed the features are public and given exogenously, a setting known as label differential privacy. In some scenarios, this can be a strong assumption that removes the interplay between features and labels from the privacy analysis. We relax this approach and instead assume the features are drawn from a distribution that depends on the private labels. We first show that simply adding noise to the label, as in previous work, can lead to an arbitrarily weak privacy guarantee, and also present methods for estimating this privacy loss from data. We then present a new mechanism that replaces some training examples with synthetically generated data, and show that our mechanism has a much better privacy-utility tradeoff if the synthetic data is realistic, in a certain quantifiable sense. Finally, we empirically validate our theoretical analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Auditing Privacy Mechanisms via Label Inference AttacksRóbert Busa-Fekete, Travis Dick, Claudio Gentile, Andrés Muñoz Medina 等NeurIPS 2024 · 被引用 3 次
- Differentially Private Analysis for Binary Response Models: Optimality, Estimation, and InferenceCe Zhang, Yixin Han, Yafei Wang, Xiaodong Yan 等ICML 2025
- Private Learning with Public Feature ConditioningShuli Jiang, Walid Krichene, Nicolas MayorazICML 2026
- Private Direct Preference Optimization for LLM AlignmentYangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao 等CCS 2026
- Enhancing Learning with Label Differential Privacy by Vector ApproximationPuning Zhao, Jiafei Wu, Zhe Liu, Li Shen 等ICLR 2025
它引用的顶会 Paper4
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Deep Learning with Label Differential PrivacyBadih Ghazi, Noah Golowich, Ravi Kumar, Pasin Manurangsi 等NeurIPS 2021 · 被引用 193 次
- Bayesian Differential Privacy for Machine LearningAleksei Triastcyn, Boi FaltingsICML 2020 · 被引用 79 次
- Lessons from the AdKDD'21 Privacy-Preserving ML ChallengeEustache Diemert, Romain Fabre, Alexandre Gilotte, Fei Jia 等WWW 2022 · 被引用 8 次
相关 Paper
- Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic DataYvonne Zhou, Mingyu Liang, Ivan Brugere, Danial Dervovic 等ICML 2024 · 被引用 3 次
- Does Training with Synthetic Data Truly Protect Privacy?Yunpeng Zhao, Jie ZhangICLR 2025
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 被引用 586 次
- Antipodes of Label Differential Privacy: PATE and ALIBIMani Malek Esmaeili, Ilya Mironov, Karthik Prasad, Igor Shilov 等NeurIPS 2021 · 被引用 84 次
- Optimal Unbiased Randomizers for Regression with Label Differential PrivacyAshwinkumar Badanidiyuru Varadaraja, Badih Ghazi, Pritish Kamath, Ravi Kumar 等NeurIPS 2023 · 被引用 9 次
