Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data
Samarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Richard Chen, Leonid Karlinsky, Kate Saenko, Venkatesh Saligrama, Rogério Schmidt Feris
Abstract
Pre-training models on Imagenet or other massive datasets of real images has led to major advances in computer vision, albeit accompanied with shortcomings related to curation cost, privacy, usage rights, and ethical issues. In this paper, for the first time, we study the transferability of pre-trained models based on synthetic data generated by graphics simulators to downstream tasks from very different domains. In using such synthetic data for pre-training, we find that downstream performance on different tasks are favored by different configurations of simulation parameters (e.g. lighting, object pose, backgrounds, etc.), and that there is no one-size-fits-all solution. It is thus better to tailor synthetic pre-training data to a specific downstream task, for best performance. We introduce Task2Sim, a unified model mapping downstream task representations to optimal simulation parameters to generate synthetic pre-training data for them. Task2Sim learns this mapping by training to find the set of best parameters on a set of "seen" tasks. Once trained, it can then be used to predict best simulation parameters for novel "unseen" tasks in one shot, without requiring additional training. Given a budget in number of images per class, our extensive experiments with 20 diverse downstream tasks show Task2Sim's task-adaptive pretraining data results in significantly better downstream performance than non-adaptively choosing simulation parameters on both seen and unseen tasks. It is even competitive with pre-training on real images from Imagenet. † Work done as interns at MIT-IBM Watson AI Lab. * Now affiliated with JPMorgan Chase, FLARE. Work done when Chun-Fu was at MIT-IBM Watson AI Lab. Project page : https://samarth4149.github.io/projects/task2sim.html Synthetic Data Pre-training Downstream Tasks from Various Domains (Real Images) ChestX Sketch Flowers SVHN EuroSAT ...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee6d170f-0193-4310-8011-92084fd141b9Cited by top-tier papers6
- Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain GeneralizationPrithvijit Chattopadhyay, Kartik Sarangmath, Vivek Vijaykumar, Judy HoffmanICCV 2023 · 57 citations
- Going Beyond Nouns With Vision & Language Models Using Synthetic DataPaola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh et al.ICCV 2023 · 49 citations
- ExPT: Synthetic Pretraining for Few-Shot Experimental DesignTung Nguyen, Sudhanshu Agrawal, Aditya GroverNeurIPS 2023 · 27 citations
- Progressive Transformation Learning for Leveraging Virtual Images in TrainingYi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. BhattacharyyaCVPR 2023
- GMAIL: Generative Modality Alignment for generated Image LearningShentong Mo, Sukmin YunICML 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization Without Accessing Target Domain DataXiangyu Yue, Yang Zhang, Sicheng Zhao, Alberto L. Sangiovanni-Vincentelli et al.ICCV 2019 · 462 citations
Related papers
- Insights into Pre-training via Simpler Synthetic TasksYuhuai Wu, Felix Li, Percy LiangNeurIPS 2022 · 29 citations
- Is Synthetic Data from Generative Models Ready for Image Recognition?Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue et al.ICLR 2023 · 56 citations
- Task-Robust Pre-Training for Worst-Case Downstream AdaptationJianghui Wang, Yang Chen, Xingyu Xie, Cong Fang et al.NeurIPS 2023 · 3 citations
- A New Benchmark: On the Utility of Synthetic Data with Blender for Bare Supervised Learning and Downstream Domain AdaptationHui Tang, Kui JiaCVPR 2023
- Automated Synthetic-to-Real GeneralizationWuyang Chen, Zhiding Yu, Zhangyang Wang, Animashree AnandkumarICML 2020 · 72 citations
