Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic Data
Boris van Breugel, Zhaozhi Qian, Mihaela van der Schaar
摘要
Generating synthetic data through generative models is gaining interest in the ML community and beyond, promising a future where datasets can be tailored to individual needs. Unfortunately, synthetic data is usually not perfect, resulting in potential errors in downstream tasks. In this work we explore how the generative process affects the downstream ML task. We show that the naive synthetic data approach -- using synthetic data as if it is real -- leads to downstream models and analyses that do not generalize well to real data. As a first step towards better ML in the synthetic data regime, we introduce Deep Generative Ensemble (DGE) -- a framework inspired by Deep Ensembles that aims to implicitly approximate the posterior distribution over the generative process model parameters. DGE improves downstream model training, evaluation, and uncertainty quantification, vastly outperforming the naive approach on average. The largest improvements are achieved for minority classes and low-density regions of the original data, for which the generative uncertainty is largest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- On the Constrained Time-Series Generation ProblemAndrea Coletta, Sriram Gopalakrishnan, Daniel Borrajo, Svitlana VyetrenkoNeurIPS 2023 · 被引用 92 次
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 被引用 51 次
- TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based ModelsAndrei Margeloiu, Xiangjian Jiang, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 被引用 19 次
- Active Learning with LLMs for Partially Observed and Cost-Aware ScenariosNicolás Astorga, Tennison Liu, Nabeel Seedat, Mihaela van der SchaarNeurIPS 2024 · 被引用 11 次
- Access Denied: Meaningful Data Access for Quantitative Algorithm AuditsJuliette Zaccour, Reuben Binns, Luc RocherCHI 2025 · 被引用 9 次
它引用的顶会 Paper6
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 被引用 287 次
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 被引用 263 次
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 被引用 174 次
- Uncertainty-Aware Deep Classifiers Using Generative ModelsMurat Sensoy, Lance M. Kaplan, Federico Cerutti, Maryam SalekiAAAI 2020 · 被引用 88 次
相关 Paper
- Debiasing Synthetic Data Generated by Deep Generative ModelsAlexander Decruyenaere, Heidelinde Dehaene, Paloma Rabaey, Johan Decruyenaere 等NeurIPS 2024 · 被引用 5 次
- High-dimensional Analysis of Synthetic Data SelectionParham Rezaei, Filip Kovacevic, Francesco Locatello, Marco MondelliICLR 2026 · 被引用 6 次
- Uncertainty Quantification with the Empirical Neural Tangent KernelJoseph Wilson, Chris van der Heide, Liam Hodgkinson, Fred RoostaNeurIPS 2025 · 被引用 11 次
- Regularizing Neural Networks with Meta-Learning Generative ModelsShin'ya Yamaguchi, Daiki Chijiwa, Sekitoshi Kanai, Atsutoshi Kumagai 等NeurIPS 2023 · 被引用 10 次
- Deep Ensembles for Graphs with Higher-order DependenciesSteven J. Krieg, William C. Burgis, Patrick M. Soga, Nitesh V. ChawlaICLR 2023 · 被引用 1 次
