Privacy Amplification Through Synthetic Data: Insights from Linear Regression
Clément Pierquin, Aurélien Bellet, Marc Tommasi, Matthieu Boussard
Abstract
Synthetic data inherits the differential privacy guarantees of the model used to generate it. Additionally, synthetic data may benefit from privacy amplification when the generative model is kept hidden. While empirical studies suggest this phenomenon, a rigorous theoretical understanding is still lacking. In this paper, we investigate this question through the well-understood framework of linear regression. First, we establish negative results showing that if an adversary controls the seed of the generative model, a single synthetic data point can leak as much information as releasing the model itself. Conversely, we show that when synthetic data is generated from random inputs, releasing a limited number of synthetic data points amplifies privacy beyond the model's inherent guarantees. We believe our findings in linear regression can serve as a foundation for deriving more general bounds in the future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46ed405c-a661-4edb-b845-0d662cca1210Cited by top-tier papers2
- Fully Decentralized Certified UnlearningHithem Lamri, Michail ManiatakosCVPR 2026 · 1 citation
- SelPE: Progressive Selection for Private Structured Text SynthesisXuancheng Zhu, Guoshun Nan, Han Zhang, Ben Niu et al.KDD 2026
Builds on8
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Deep Learning with Label Differential PrivacyBadih Ghazi, Noah Golowich, Ravi Kumar, Pasin Manurangsi et al.NeurIPS 2021 · 193 citations
- Model Collapse Demystified: The Case of RegressionElvis Dohmatob, Yunzhen Feng, Julia KempeNeurIPS 2024 · 96 citations
- Privacy of Noisy Stochastic Gradient Descent: More Iterations without More Privacy LossJason M. Altschuler, Kunal TalwarNeurIPS 2022 · 89 citations
- SoK: Privacy-Preserving Data SynthesisYuzheng Hu, Fan Wu, Qinbin Li, Yunhui Long et al.S&P 2024 · 61 citations
Related papers
- Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic DataYvonne Zhou, Mingyu Liang, Ivan Brugere, Danial Dervovic et al.ICML 2024 · 3 citations
- Synthetic Data - Anonymisation Groundhog DayTheresa Stadler, Bristena Oprisanu, Carmela TroncosoUSENIX Security 2022
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 37 citations
- Does Training with Synthetic Data Truly Protect Privacy?Yunpeng Zhao, Jie ZhangICLR 2025
- Label differential privacy and private training data releaseRóbert Istvan Busa-Fekete, Andrés Muñoz Medina, Umar Syed, Sergei VassilvitskiiICML 2023 · 9 citations
