Does Training with Synthetic Data Truly Protect Privacy?
Yunpeng Zhao, Jie Zhang
摘要
As synthetic data becomes increasingly popular in machine learning tasks, numerous methods-without formal differential privacy guarantees-use synthetic data for training. These methods often claim, either explicitly or implicitly, to protect the privacy of the original training data. In this work, we explore four different training paradigms: coreset selection, dataset distillation, data-free knowledge distillation, and synthetic data generated from diffusion models. While all these methods utilize synthetic data for training, they lead to vastly different conclusions regarding privacy preservation. We caution that empirical approaches to preserving data privacy require careful and rigorous evaluation; otherwise, they risk providing a false sense of privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- FedGPS: Statistical Rectification Against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Chenxin Li, Yiu-ming Cheung 等NeurIPS 2025 · 被引用 7 次
- CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset TrainingQi Li, Cheng-Long Wang, Yinzhi Cao, Di WangACL 2026 · 被引用 2 次
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 被引用 2 次
- Learnability and Privacy Vulnerability are Entangled in a Few Critical WeightsXingli Fang, Jung-Eun KimICLR 2026 · 被引用 1 次
- CompLeak: Deep Learning Model Compression Exacerbates Privacy LeakageNa Li, Yansong Gao, Hongsheng Hu, Boyu Kuang 等USENIX Security 2026
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
相关 Paper
- Label differential privacy and private training data releaseRóbert Istvan Busa-Fekete, Andrés Muñoz Medina, Umar Syed, Sergei VassilvitskiiICML 2023 · 被引用 9 次
- Students Parrot Their Teachers: Membership Inference on Model DistillationMatthew Jagielski, Milad Nasr, Katherine Lee, Christopher A. Choquette-Choo 等NeurIPS 2023 · 被引用 53 次
- Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic DataYvonne Zhou, Mingyu Liang, Ivan Brugere, Danial Dervovic 等ICML 2024 · 被引用 3 次
- dp-promise: Differentially Private Diffusion Probabilistic Models for Image SynthesisHaichen Wang, Shuchao Pang, Zhigang Lu, Yihang Rao 等USENIX Security 2024 · 被引用 36 次
- Privacy Amplification Through Synthetic Data: Insights from Linear RegressionClément Pierquin, Aurélien Bellet, Marc Tommasi, Matthieu BoussardICML 2025
