Private Synthetic Data for Multitask Learning and Marginal Queries
Giuseppe Vietri, Cédric Archambeau, Sergül Aydöre, William Brown, Michael Kearns, Aaron Roth, Amaresh Ankit Siva, Shuai Tang, Zhiwei Steven Wu
摘要
We provide a differentially private algorithm for producing synthetic data simultaneously useful for multiple tasks: marginal queries and multitask machine learning (ML). A key innovation in our algorithm is the ability to directly handle numerical features, in contrast to a number of related prior approaches which require numerical features to be first converted into high cardinality categorical features via a binning strategy. Higher binning granularity is required for better accuracy, but this negatively impacts scalability. Eliminating the need for binning allows us to produce synthetic data preserving large numbers of statistical queries such as marginals on numerical features, and class conditional linear threshold queries. Preserving the latter means that the fraction of points of each class label above a particular half-space is roughly the same in both the real and synthetic data. This is the property that is needed to train a linear classifier in a multitask setting. Our algorithm also allows us to produce high quality synthetic data for mixed marginal queries, that combine both categorical and numerical features. Our method consistently runs 2-5x faster than the best comparable techniques, and provides significant accuracy improvements in both marginal queries and linear prediction tasks for mixed-type datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 被引用 41 次
- Generating Private Synthetic Data with Genetic AlgorithmsTerrance Liu, Jingwu Tang, Giuseppe Vietri, Steven WuICML 2023 · 被引用 20 次
- Post-processing Private Synthetic Data for Improving Utility on Selected MeasuresHao Wang, Shivchander Sudalairaj, John Henning, Kristjan H. Greenewald 等NeurIPS 2023 · 被引用 13 次
- CuTS: Customizable Tabular Synthetic Data GenerationMark Vero, Mislav Balunovic, Martin T. VechevICML 2024 · 被引用 13 次
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential PrivacyVishnu Vinod, Krishna Pillutla, Abhradeep Guha ThakurtaNeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper8
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
- New Oracle-Efficient Algorithms for Private Synthetic Data ReleaseGiuseppe Vietri, Grace Tian, Mark Bun, Thomas Steinke 等ICML 2020 · 被引用 86 次
- Iterative Methods for Private Synthetic Data: Unifying Framework and New MethodsTerrance Liu, Giuseppe Vietri, Steven WuNeurIPS 2021 · 被引用 85 次
- Differentially Private Query Release Through Adaptive ProjectionSergül Aydöre, William Brown, Michael Kearns, Krishnaram Kenthapadi 等ICML 2021 · 被引用 78 次
相关 Paper
- The Importance of Being Discrete: Measuring the Impact of Discretization in End-to-End Differentially Private Synthetic DataGeorgi Ganev, Meenatchi Sundaram Muthu Selva Annamalai, Sofiane Mahiou, Emiliano De CristofaroCCS 2025
- Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic DataYvonne Zhou, Mingyu Liang, Ivan Brugere, Danial Dervovic 等ICML 2024 · 被引用 3 次
- Privacy-Preserving Data Release Leveraging Optimal Transport and Particle Gradient DescentKonstantin Donhauser, Javier Abad Martinez, Neha Hulkund, Fanny YangICML 2024 · 被引用 6 次
- Differentially Private Sum-Product NetworksXenia Heilmann, Mattia Cerrato, Ernst AlthausICML 2024 · 被引用 1 次
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 被引用 37 次
