Provable Privacy with Non-Private Pre-Processing
Yaxi Hu, Amartya Sanyal, Bernhard Schölkopf
Abstract
When analyzing Differentially Private (DP) machine learning pipelines, the potential privacy cost of data-dependent pre-processing is frequently overlooked in privacy accounting. In this work, we propose a general framework to evaluate the additional privacy cost incurred by non-private data-dependent pre-processing algorithms. Our framework establishes upper bounds on the overall privacy guarantees by utilising two new technical notions: a variant of DP termed Smooth DP and the bounded sensitivity of the pre-processing algorithms. In addition to the generic framework, we provide explicit overall privacy guarantees for multiple data-dependent pre-processing algorithms, such as data imputation, quantization, deduplication, standard scaling and PCA, when used in combination with several DP algorithms. Notably, this framework is also simple to implement, allowing direct integration into existing DP pipelines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 290d4ffe-47e2-4eb1-a705-36a8de6749f3Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
Related papers
- Individual Sensitivity Preprocessing for Data PrivacyRachel Cummings, David DurfeeSODA 2020 · 29 citations
- Machine Learning with Privacy for Protected AttributesSaeed Mahloujifar, Chuan Guo, G. Edward Suh, Kamalika ChaudhuriS&P 2025
- Unified Mechanism-Specific Amplification by Subsampling and Group Privacy AmplificationJan Schuchardt, Mihail Stoian, Arthur Kosmala, Stephan GünnemannNeurIPS 2024 · 8 citations
- Bayesian Differential Privacy for Machine LearningAleksei Triastcyn, Boi FaltingsICML 2020 · 79 citations
- PEARL: Data Synthesis via Private Embeddings and Adversarial Reconstruction LearningSeng Pei Liew, Tsubasa Takahashi, Michihiko UenoICLR 2022 · 32 citations
