Privacy for Free: How does Dataset Condensation Help Privacy?
Tian Dong, Bo Zhao, Lingjuan Lyu
Abstract
To prevent unintentional data leakage, research community has resorted to data generators that can produce differentially private data for model training. However, for the sake of the data privacy, existing solutions suffer from either expensive training cost or poor generalization performance. Therefore, we raise the question whether training efficiency and privacy can be achieved simultaneously. In this work, we for the first time identify that dataset condensation (DC) which is originally designed for improving training efficiency is also a better solution to replace the traditional data generators for private data generation, thus providing privacy for free. To demonstrate the privacy benefit of DC, we build a connection between DC and differential privacy, and theoretically prove on linear feature extractors (and then extended to non-linear feature extractors) that the existence of one sample has limited impact () on the parameter distribution of networks trained on samples synthesized from raw samples by DC. We also empirically validate the visual privacy and membership privacy of DC-synthesized data by launching both the loss-based and the state-of-the-art likelihood-based membership inference attacks. We envision this work as a milestone for data-efficient and privacy-preserving machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers49
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- DataDAM: Efficient Dataset Distillation with Attention MatchingAhmad Sajedi, Samir Khaki, Ehsan Amjadian, Lucy Z. Liu et al.ICCV 2023 · 106 citations
- Condensing Graphs via One-Step Gradient MatchingWei Jin, Xianfeng Tang, Haoming Jiang, Zheng Li et al.KDD 2022 · 68 citations
- Dataset Distillation with Convexified Implicit GradientsNoel Loo, Ramin M. Hasani, Mathias Lechner, Daniela RusICML 2023 · 56 citations
- Private Set Generation with Discriminative InformationDingfan Chen, Raouf Kerkouche, Mario FritzNeurIPS 2022 · 51 citations
Builds on19
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 1,736 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 628 citations
Related papers
- DP-GenG: Differentially Private Dataset Distillation Guided by DP-Generated DataShuo Shi, Jinghuai Zhang, Shijie Jiang, Chunyi Zhou et al.AAAI 2026
- RelaxLoss: Defending Membership Inference Attacks without Losing UtilityDingfan Chen, Ning Yu, Mario FritzICLR 2022 · 61 citations
- Students Parrot Their Teachers: Membership Inference on Model DistillationMatthew Jagielski, Milad Nasr, Katherine Lee, Christopher A. Choquette-Choo et al.NeurIPS 2023 · 53 citations
- Improving Noise Efficiency in Privacy-Preserving Dataset DistillationRunkai Zheng, Vishnu Asutosh Dasu, Yinong Oliver Wang, Haohan Wang et al.ICCV 2025
- PAR-GAN: Improving the Generalization of Generative Adversarial Networks Against Membership Inference AttacksJunjie Chen, Wendy Hui Wang, Hongchang Gao, Xinghua ShiKDD 2021 · 29 citations
