Synthetic data for model selection
Alon Shoshan, Nadav Bhonker, Igor Kviatkovsky, Matan Fintz, Gérard G. Medioni
摘要
Recent breakthroughs in synthetic data generation approaches made it possible to produce highly photorealistic images which are hardly distinguishable from real ones. Furthermore, synthetic generation pipelines have the potential to generate an unlimited number of images. The combination of high photorealism and scale turn synthetic data into a promising candidate for improving various machine learning (ML) pipelines. Thus far, a large body of research in this field has focused on using synthetic images for training, by augmenting and enlarging training data. In contrast to using synthetic data for training, in this work we explore whether synthetic data can be beneficial for model selection. Considering the task of image classification, we demonstrate that when data is scarce, synthetic data can be used to replace the held out validation set, thus allowing to train on a larger dataset. We also introduce a novel method to calibrate the synthetic error estimation to fit that of the real domain. We show that such calibration significantly improves the usefulness of synthetic data for model selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fill In The Gaps: Model Calibration and Generalization with Synthetic DataYang Ba, Michelle Mancenido, Rong PanEMNLP 2024 · 被引用 2 次
- When Sample Selection Bias Precipitates Model CollapseXinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang 等ICML 2026
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Is Synthetic Data from Generative Models Ready for Image Recognition?Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue 等ICLR 2023 · 被引用 56 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
- AUGCAL: Improving Sim2Real Adaptation by Uncertainty Calibration on Augmented Synthetic ImagesPrithvijit Chattopadhyay, Bharat Goyal, Boglarka Ecsedi, Viraj Prabhu 等ICLR 2024 · 被引用 2 次
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian 等NeurIPS 2022 · 被引用 69 次
- When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data GeneratorsKrzysztof Adamkiewicz, Brian B. Moser, Stanislav Frolov, Tobias Christian Nauen 等CVPR 2026 · 被引用 8 次
