Label differential privacy and private training data release
Róbert Istvan Busa-Fekete, Andrés Muñoz Medina, Umar Syed, Sergei Vassilvitskii
Abstract
We study differentially private mechanisms for sharing training data in machine learning settings. Our goal is to enable learning of an accurate predictive model while protecting the privacy of each user's label. Previous work established privacy guarantees that assumed the features are public and given exogenously, a setting known as label differential privacy. In some scenarios, this can be a strong assumption that removes the interplay between features and labels from the privacy analysis. We relax this approach and instead assume the features are drawn from a distribution that depends on the private labels. We first show that simply adding noise to the label, as in previous work, can lead to an arbitrarily weak privacy guarantee, and also present methods for estimating this privacy loss from data. We then present a new mechanism that replaces some training examples with synthetically generated data, and show that our mechanism has a much better privacy-utility tradeoff if the synthetic data is realistic, in a certain quantifiable sense. Finally, we empirically validate our theoretical analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd7272a1-52ba-4125-ac76-8a999f521318Cited by top-tier papers6
- Auditing Privacy Mechanisms via Label Inference AttacksRóbert Busa-Fekete, Travis Dick, Claudio Gentile, Andrés Muñoz Medina et al.NeurIPS 2024 · 3 citations
- Differentially Private Analysis for Binary Response Models: Optimality, Estimation, and InferenceCe Zhang, Yixin Han, Yafei Wang, Xiaodong Yan et al.ICML 2025
- Private Learning with Public Feature ConditioningShuli Jiang, Walid Krichene, Nicolas MayorazICML 2026
- Private Direct Preference Optimization for LLM AlignmentYangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao et al.CCS 2026
- Enhancing Learning with Label Differential Privacy by Vector ApproximationPuning Zhao, Jiafei Wu, Zhe Liu, Li Shen et al.ICLR 2025
Builds on4
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Deep Learning with Label Differential PrivacyBadih Ghazi, Noah Golowich, Ravi Kumar, Pasin Manurangsi et al.NeurIPS 2021 · 193 citations
- Bayesian Differential Privacy for Machine LearningAleksei Triastcyn, Boi FaltingsICML 2020 · 79 citations
- Lessons from the AdKDD'21 Privacy-Preserving ML ChallengeEustache Diemert, Romain Fabre, Alexandre Gilotte, Fei Jia et al.WWW 2022 · 8 citations
Related papers
- Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic DataYvonne Zhou, Mingyu Liang, Ivan Brugere, Danial Dervovic et al.ICML 2024 · 3 citations
- Does Training with Synthetic Data Truly Protect Privacy?Yunpeng Zhao, Jie ZhangICLR 2025
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 586 citations
- Antipodes of Label Differential Privacy: PATE and ALIBIMani Malek Esmaeili, Ilya Mironov, Karthik Prasad, Igor Shilov et al.NeurIPS 2021 · 84 citations
- Optimal Unbiased Randomizers for Regression with Label Differential PrivacyAshwinkumar Badanidiyuru Varadaraja, Badih Ghazi, Pritish Kamath, Ravi Kumar et al.NeurIPS 2023 · 9 citations
