Learning from Mixtures of Private and Public Populations
Raef Bassily, Shay Moran, Anupama Nandi
Abstract
We initiate the study of a new model of supervised learning under privacy constraints. Imagine a medical study where a dataset is sampled from a population of both healthy and unhealthy individuals. Suppose healthy individuals have no privacy concerns (in such case, we call their data "public") while the unhealthy individuals desire stringent privacy protection for their data. In this example, the population (data distribution) is a mixture of private (unhealthy) and public (healthy) sub-populations that could be very different. Inspired by the above example, we consider a model in which the population is a mixture of two sub-populations: a private sub-population of private and sensitive data, and a public sub-population of data with no privacy concerns. Each example drawn from is assumed to contain a privacy-status bit that indicates whether the example is private or public. The goal is to design a learning algorithm that satisfies differential privacy only with respect to the private examples. Prior works in this context assumed a homogeneous population where private and public data arise from the same distribution, and in particular designed solutions which exploit this assumption. We demonstrate how to circumvent this assumption by considering, as a case study, the problem of learning linear classifiers in . We show that in the case where the privacy status is correlated with the target label (as in the above example), linear classifiers in can be learned, in the agnostic as well as the realizable setting, with sample complexity which is comparable to that of the classical (non-private) PAC-learning. It is known that this task is impossible if all the data is considered private.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- Iterative Methods for Private Synthetic Data: Unifying Framework and New MethodsTerrance Liu, Giuseppe Vietri, Steven WuNeurIPS 2021 · 85 citations
- Leveraging Public Data for Practical Private Query ReleaseTerrance Liu, Giuseppe Vietri, Thomas Steinke, Jonathan R. Ullman et al.ICML 2021 · 68 citations
- Public Data-Assisted Mirror Descent for Private Model TrainingEhsan Amid, Arun Ganesh, Rajiv Mathews, Swaroop Ramaswamy et al.ICML 2022 · 61 citations
- Private Estimation with Public DataAlex Bie, Gautam Kamath, Vikrant SinghalNeurIPS 2022 · 40 citations
Builds on1
Related papers
- Oracle-Efficient Differentially Private Learning with Public DataAdam Block, Mark Bun, Rathin Desai, Abhishek Shetty et al.NeurIPS 2024 · 6 citations
- Label differential privacy and private training data releaseRóbert Istvan Busa-Fekete, Andrés Muñoz Medina, Umar Syed, Sergei VassilvitskiiICML 2023 · 9 citations
- Private Distribution Learning with Public Data: The View from Sample CompressionShai Ben-David, Alex Bie, Clément L. Canonne, Gautam Kamath et al.NeurIPS 2023 · 18 citations
- Differentially Private Domain Adaptation with Theoretical GuaranteesRaef Bassily, Corinna Cortes, Anqi Mao, Mehryar MohriICML 2024
- Littlestone Classes are Privately Online LearnableNoah Golowich, Roi LivniNeurIPS 2021 · 15 citations
