Post-processing Private Synthetic Data for Improving Utility on Selected Measures
Hao Wang, Shivchander Sudalairaj, John Henning, Kristjan H. Greenewald, Akash Srivastava
Abstract
Existing private synthetic data generation algorithms are agnostic to downstream tasks. However, end users may have specific requirements that the synthetic data must satisfy. Failure to meet these requirements could significantly reduce the utility of the data for downstream use. We introduce a post-processing technique that improves the utility of the synthetic data with respect to measures selected by the end user, while preserving strong privacy guarantees and dataset quality. Our technique involves resampling from the synthetic data to filter out samples that do not meet the selected utility measures, using an efficient stochastic first-order algorithm to find optimal resampling weights. Through comprehensive numerical experiments, we demonstrate that our approach consistently improves the utility of synthetic data across multiple benchmark datasets and state-of-the-art synthetic data generation algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d9abbf0-5ca5-4e59-bcda-491532b6a81dCited by top-tier papers3
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 41 citations
- Privacy without Noisy Gradients: Slicing Mechanism for Generative Model TrainingKristjan H. Greenewald, Yuancheng Yu, Hao Wang, Kai XuNeurIPS 2024 · 5 citations
- Optimal Domain-Aware Privacy Mechanisms for Synthetic Data GenerationSajani Vithana, Sangwon Jung, Haoyang Hu, Viveck Cadambe et al.ICML 2026
Builds on10
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 136 citations
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale et al.NeurIPS 2021 · 113 citations
- Iterative Methods for Private Synthetic Data: Unifying Framework and New MethodsTerrance Liu, Giuseppe Vietri, Steven WuNeurIPS 2021 · 85 citations
- Differentially Private Query Release Through Adaptive ProjectionSergül Aydöre, William Brown, Michael Kearns, Krishnaram Kenthapadi et al.ICML 2021 · 78 citations
- Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic DataGeorgi Ganev, Bristena Oprisanu, Emiliano De CristofaroICML 2022 · 78 citations
Related papers
- Private Post-GAN BoostingMarcel Neunhoeffer, Steven Wu, Cynthia DworkICLR 2021 · 30 citations
- Private Set Generation with Discriminative InformationDingfan Chen, Raouf Kerkouche, Mario FritzNeurIPS 2022 · 51 citations
- Differentially Private Data Generation with Missing DataShubhankar Mohapatra, Jianqiao Zong, Florian Kerschbaum, Xi HeVLDB 2024 · 7 citations
- PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic DataXinyuan Zhao, Hanlin Gu, Guibao Song, Gongxi Zhu et al.CVPR 2026
- Epistemic Parity: Reproducibility as an Evaluation Metric for Differential PrivacyLucas Rosenblatt, Bernease Herman, Anastasia Holovenko, Wonkwon Lee et al.VLDB 2023 · 11 citations
