High-dimensional Analysis of Synthetic Data Selection
Parham Rezaei, Filip Kovacevic, Francesco Locatello, Marco Mondelli
摘要
Despite the progress in the development of generative models, their usefulness in creating synthetic data that improve prediction performance of classifiers has been put into question. Besides heuristic principles such as "synthetic data should be close to the real data distribution", it is actually not clear which specific properties affect the generalization error. Our paper addresses this question through the lens of high-dimensional regression. Theoretically, we show that, for linear models, the covariance shift between the target distribution and the distribution of the synthetic data affects the generalization error but, surprisingly, the mean shift does not. Furthermore we prove that, in some settings, matching the covariance of the target distribution is optimal. Remarkably, the theoretical insights from linear models carry over to deep neural networks and generative models. We empirically demonstrate that the covariance matching procedure (matching the covariance of the synthetic data with that of the data coming from the target distribution) performs well against several recent approaches for synthetic data selection, across training paradigms, architectures, datasets and generative models used for augmentation. 1 How to select the dataset (X s , y s ) in order to minimize the test error? (Q) By studying this question, we can identify which properties of the distribution of (X s , y s ) improve generalization, thus guiding the selection of data obtained in practice from generative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 被引用 9 次
- Optimal Regularization for Performative LearningEdwige Cyffers, Alireza Mirrokni, Marco MondelliICML 2026 · 被引用 1 次
- When Sample Selection Bias Precipitates Model CollapseXinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang 等ICML 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi 等ICML 2020 · 被引用 553 次
相关 Paper
- Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic DataBoris van Breugel, Zhaozhi Qian, Mihaela van der SchaarICML 2023 · 被引用 45 次
- Selective Mixup Helps with Distribution Shifts, But Not (Only) because of MixupDamien Teney, Jindong Wang, Ehsan AbbasnejadICML 2024 · 被引用 9 次
- On Predicting Generalization using GANsYi Zhang, Arushi Gupta, Nikunj Saunshi, Sanjeev AroraICLR 2022 · 被引用 8 次
- Regularizing Neural Networks with Meta-Learning Generative ModelsShin'ya Yamaguchi, Daiki Chijiwa, Sekitoshi Kanai, Atsutoshi Kumagai 等NeurIPS 2023 · 被引用 10 次
- AEC-GAN: Adversarial Error Correction GANs for Auto-Regressive Long Time-Series GenerationLei Wang, Liang Zeng, Jian LiAAAI 2023 · 被引用 16 次
