Insights into Pre-training via Simpler Synthetic Tasks
Yuhuai Wu, Felix Li, Percy Liang
Abstract
Pre-training produces representations that are effective for a wide range of downstream tasks, but it is still unclear what properties of pre-training are necessary for effective gains. Notably, recent work shows that even pre-training on synthetic tasks can achieve significant gains in downstream tasks. In this work, we perform three experiments that iteratively simplify pre-training and show that the simplifications still retain much of its gains. First, building on prior work, we perform a systematic evaluation of three existing synthetic pre-training methods on six downstream tasks. We find the best synthetic pre-training method, LIME, attains an average of of the benefits of natural pre-training. Second, to our surprise, we find that pre-training on a simple and generic synthetic task defined by the Set function achieves of the benefits, almost matching LIME. Third, we find that of the benefits can be attained by using merely the parameter statistics of synthetic pre-training. We release the source code at https://github.com/felixzli/synthetic_pretraining.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 337eeba1-e2c4-48d0-af4d-542fa893a977Cited by top-tier papers9
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng et al.NeurIPS 2023 · 149 citations
- Downstream Datasets Make Surprisingly Good Pretraining CorporaKundan Krishna, Saurabh Garg, Jeffrey P. Bigham, Zachary C. LiptonACL 2023 · 11 citations
- Pre-training with Synthetic Data Helps Offline Reinforcement LearningZecheng Wang, Che Wang, Zixuan Dong, Keith W. RossICLR 2024 · 11 citations
- On the Importance and Applicability of Pre-Training for Federated LearningHong-You Chen, Cheng-Hao Tu, Ziwei Li, Han-Wei Shen et al.ICLR 2023 · 10 citations
- Can You Learn to See Without Images? Procedural Warm-Up for Vision TransformersZachary Shinnick, Liangze Jiang, Hemanth Saratchandran, Damien Teney et al.CVPR 2026 · 9 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong et al.EMNLP 2022 · 222 citations
Related papers
- How Useful Is Self-Supervised Pretraining for Visual Tasks?Alejandro Newell, Jia DengCVPR 2020
- Task2Sim: Towards Effective Pre-training and Transfer from Synthetic DataSamarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Richard Chen et al.CVPR 2022 · 28 citations
- LIME: Learning Inductive Bias for Primitives of Mathematical ReasoningYuhuai Wu, Markus N. Rabe, Wenda Li, Jimmy Ba et al.ICML 2021 · 66 citations
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Task-Robust Pre-Training for Worst-Case Downstream AdaptationJianghui Wang, Yang Chen, Xingyu Xie, Cong Fang et al.NeurIPS 2023 · 3 citations
