High-dimensional Analysis of Synthetic Data Selection
Parham Rezaei, Filip Kovacevic, Francesco Locatello, Marco Mondelli
Abstract
Despite the progress in the development of generative models, their usefulness in creating synthetic data that improve prediction performance of classifiers has been put into question. Besides heuristic principles such as "synthetic data should be close to the real data distribution", it is actually not clear which specific properties affect the generalization error. Our paper addresses this question through the lens of high-dimensional regression. Theoretically, we show that, for linear models, the covariance shift between the target distribution and the distribution of the synthetic data affects the generalization error but, surprisingly, the mean shift does not. Furthermore we prove that, in some settings, matching the covariance of the target distribution is optimal. Remarkably, the theoretical insights from linear models carry over to deep neural networks and generative models. We empirically demonstrate that the covariance matching procedure (matching the covariance of the synthetic data with that of the data coming from the target distribution) performs well against several recent approaches for synthetic data selection, across training paradigms, architectures, datasets and generative models used for augmentation. 1 How to select the dataset (X s , y s ) in order to minimize the test error? (Q) By studying this question, we can identify which properties of the distribution of (X s , y s ) improve generalization, thus guiding the selection of data obtained in practice from generative models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecfef1c3-ef46-4f6c-9a9d-cabfb1ac9cfeCited by top-tier papers3
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 9 citations
- Optimal Regularization for Performative LearningEdwige Cyffers, Alireza Mirrokni, Marco MondelliICML 2026 · 1 citation
- When Sample Selection Bias Precipitates Model CollapseXinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang et al.ICML 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi et al.ICML 2020 · 553 citations
Related papers
- Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic DataBoris van Breugel, Zhaozhi Qian, Mihaela van der SchaarICML 2023 · 45 citations
- Selective Mixup Helps with Distribution Shifts, But Not (Only) because of MixupDamien Teney, Jindong Wang, Ehsan AbbasnejadICML 2024 · 9 citations
- On Predicting Generalization using GANsYi Zhang, Arushi Gupta, Nikunj Saunshi, Sanjeev AroraICLR 2022 · 8 citations
- Regularizing Neural Networks with Meta-Learning Generative ModelsShin'ya Yamaguchi, Daiki Chijiwa, Sekitoshi Kanai, Atsutoshi Kumagai et al.NeurIPS 2023 · 10 citations
- AEC-GAN: Adversarial Error Correction GANs for Auto-Regressive Long Time-Series GenerationLei Wang, Liang Zeng, Jian LiAAAI 2023 · 16 citations
