Post-processing Private Synthetic Data for Improving Utility on Selected Measures
Hao Wang, Shivchander Sudalairaj, John Henning, Kristjan H. Greenewald, Akash Srivastava
摘要
Existing private synthetic data generation algorithms are agnostic to downstream tasks. However, end users may have specific requirements that the synthetic data must satisfy. Failure to meet these requirements could significantly reduce the utility of the data for downstream use. We introduce a post-processing technique that improves the utility of the synthetic data with respect to measures selected by the end user, while preserving strong privacy guarantees and dataset quality. Our technique involves resampling from the synthetic data to filter out samples that do not meet the selected utility measures, using an efficient stochastic first-order algorithm to find optimal resampling weights. Through comprehensive numerical experiments, we demonstrate that our approach consistently improves the utility of synthetic data across multiple benchmark datasets and state-of-the-art synthetic data generation algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 被引用 41 次
- Privacy without Noisy Gradients: Slicing Mechanism for Generative Model TrainingKristjan H. Greenewald, Yuancheng Yu, Hao Wang, Kai XuNeurIPS 2024 · 被引用 5 次
- Optimal Domain-Aware Privacy Mechanisms for Synthetic Data GenerationSajani Vithana, Sangwon Jung, Haoyang Hu, Viveck Cadambe 等ICML 2026
它引用的顶会 Paper10
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale 等NeurIPS 2021 · 被引用 113 次
- Iterative Methods for Private Synthetic Data: Unifying Framework and New MethodsTerrance Liu, Giuseppe Vietri, Steven WuNeurIPS 2021 · 被引用 85 次
- Differentially Private Query Release Through Adaptive ProjectionSergül Aydöre, William Brown, Michael Kearns, Krishnaram Kenthapadi 等ICML 2021 · 被引用 78 次
- Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic DataGeorgi Ganev, Bristena Oprisanu, Emiliano De CristofaroICML 2022 · 被引用 78 次
相关 Paper
- Private Post-GAN BoostingMarcel Neunhoeffer, Steven Wu, Cynthia DworkICLR 2021 · 被引用 30 次
- Private Set Generation with Discriminative InformationDingfan Chen, Raouf Kerkouche, Mario FritzNeurIPS 2022 · 被引用 51 次
- Differentially Private Data Generation with Missing DataShubhankar Mohapatra, Jianqiao Zong, Florian Kerschbaum, Xi HeVLDB 2024 · 被引用 7 次
- PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic DataXinyuan Zhao, Hanlin Gu, Guibao Song, Gongxi Zhu 等CVPR 2026
- Epistemic Parity: Reproducibility as an Evaluation Metric for Differential PrivacyLucas Rosenblatt, Bernease Herman, Anastasia Holovenko, Wonkwon Lee 等VLDB 2023 · 被引用 11 次
