Differentially Private Synthetic Data via Foundation Model APIs 1: Images
Zinan Lin, Sivakanth Gopi, Janardhan Kulkarni, Harsha Nori, Sergey Yekhanin
Abstract
Generating differentially private (DP) synthetic data that closely resembles the original private data is a scalable way to mitigate privacy concerns in the current data-driven world. In contrast to current practices that train customized models for this task, we aim to generate DP Synthetic Data via APIs (DPSDA), where we treat foundation models as blackboxes and only utilize their inference APIs. Such API-based, training-free approaches are easier to deploy as exemplified by the recent surge in the number of API-based apps. These approaches can also leverage the power of large foundation models which are only accessible via their inference APIs. However, this comes with greater challenges due to strictly more restrictive model access and the need to protect privacy from the API provider. In this paper, we present a new framework called Private Evolution (PE) to solve this problem and show its initial promise on synthetic images. Surprisingly, PE can match or even outperform state-of-the-art (SOTA) methods without any model training. For example, on CIFAR10 (with ImageNet as the public data), we achieve FID<= 7.9 with privacy cost = 0.67, significantly improving the previous SOTA from = 32. We further demonstrate the promise of applying PE on large foundation models such as Stable Diffusion to tackle challenging private datasets with a small number of high-resolution images. The code and data are released at https://github.com/microsoft/DPSDA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 41 citations
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMsCharlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway et al.ICML 2024 · 30 citations
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential PrivacyVishnu Vinod, Krishna Pillutla, Abhradeep Guha ThakurtaNeurIPS 2025 · 12 citations
- Synthesize Privacy-Preserving High-Resolution Images via Private Textual IntermediariesHaoxiang Wang, Zinan Lin, Da Yu, Huishuai ZhangNeurIPS 2025 · 9 citations
- Hot PATE: Private Aggregation of Distributions for Diverse TasksEdith Cohen, Benjamin Cohen-Wang, Xin Lyu, Jelani Nelson et al.ICLR 2026 · 5 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIsJianqing Zhang, Yang Liu, Jie Fu, Yang Hua et al.ICML 2025
- Private Evolution ConvergesTomás González Lara, Giulia Fanti, Aaditya RamdasNeurIPS 2025 · 3 citations
- Differentially Private Synthetic Data via Foundation Model APIs 2: TextChulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi et al.ICML 2024 · 71 citations
- Secret-Protected Evolution for Differentially Private Synthetic Text GenerationTianze Wang, Zhaoyu Chen, Jian Du, Yingtai Xiao et al.ICLR 2026 · 1 citation
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao et al.USENIX Security 2024 · 23 citations
