Private Set Generation with Discriminative Information
Dingfan Chen, Raouf Kerkouche, Mario Fritz
Abstract
Differentially private data generation techniques have become a promising solution to the data privacy challenge --it enables sharing of data while complying with rigorous privacy guarantees, which is essential for scientific progress in sensitive domains. Unfortunately, restricted by the inherent complexity of modeling highdimensional distributions, existing private generative models are struggling with the utility of synthetic samples. In contrast to existing works that aim at fitting the complete data distribution, we directly optimize for a small set of samples that are representative of the distribution under the supervision of discriminative information from downstream tasks, which is generally an easier task and more suitable for private training. Our work provides an alternative view for differentially private generation of high-dimensional data and introduces a simple yet effective method that greatly improves the sample utility of state-of-the-art approaches. The source code is available at https://github.com/DingfanChen/Private-Set .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97ca4e53-a19d-4b1a-b5a6-d4e1f879cadbCited by top-tier papers13
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- Embarrassingly Simple Dataset DistillationYunzhen Feng, Shanmukha Ramakrishna Vedantam, Julia KempeICLR 2024 · 21 citations
- Hyperbolic Dataset DistillationWenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa et al.NeurIPS 2025 · 17 citations
- DPMLBench: Holistic Evaluation of Differentially Private Machine LearningChengkun Wei, Minghu Zhao, Zhikun Zhang, Min Chen et al.CCS 2023 · 5 citations
- On the Size and Approximation Error of Distilled DatasetsAlaa Maalouf, Murad Tukan, Noel Loo, Ramin M. Hasani et al.NeurIPS 2023 · 2 citations
Builds on13
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
Related papers
- G-PATE: Scalable Differentially Private Data Generator via Private Aggregation of Teacher DiscriminatorsYunhui Long, Boxin Wang, Zhuolin Yang, Bhavya Kailkhura et al.NeurIPS 2021 · 91 citations
- DPGEN: Differentially Private Generative Energy-Guided Network for Natural Image SynthesisJia-Wei Chen, Chia-Mu Yu, Ching-Chia Kao, Tzai-Wei Pang et al.CVPR 2022 · 14 citations
- Post-processing Private Synthetic Data for Improving Utility on Selected MeasuresHao Wang, Shivchander Sudalairaj, John Henning, Kristjan H. Greenewald et al.NeurIPS 2023 · 13 citations
- Private Gradient Estimation is Useful for Generative ModelingBochao Liu, Pengju Wang, Weijia Guo, Yong Li et al.ACM MM 2024 · 1 citation
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao et al.USENIX Security 2024 · 23 citations
