Insights into Pre-training via Simpler Synthetic Tasks
Yuhuai Wu, Felix Li, Percy Liang
摘要
Pre-training produces representations that are effective for a wide range of downstream tasks, but it is still unclear what properties of pre-training are necessary for effective gains. Notably, recent work shows that even pre-training on synthetic tasks can achieve significant gains in downstream tasks. In this work, we perform three experiments that iteratively simplify pre-training and show that the simplifications still retain much of its gains. First, building on prior work, we perform a systematic evaluation of three existing synthetic pre-training methods on six downstream tasks. We find the best synthetic pre-training method, LIME, attains an average of of the benefits of natural pre-training. Second, to our surprise, we find that pre-training on a simple and generic synthetic task defined by the Set function achieves of the benefits, almost matching LIME. Third, we find that of the benefits can be attained by using merely the parameter statistics of synthetic pre-training. We release the source code at https://github.com/felixzli/synthetic_pretraining.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng 等NeurIPS 2023 · 被引用 149 次
- Downstream Datasets Make Surprisingly Good Pretraining CorporaKundan Krishna, Saurabh Garg, Jeffrey P. Bigham, Zachary C. LiptonACL 2023 · 被引用 11 次
- Pre-training with Synthetic Data Helps Offline Reinforcement LearningZecheng Wang, Che Wang, Zixuan Dong, Keith W. RossICLR 2024 · 被引用 11 次
- On the Importance and Applicability of Pre-Training for Federated LearningHong-You Chen, Cheng-Hao Tu, Ziwei Li, Han-Wei Shen 等ICLR 2023 · 被引用 10 次
- Can You Learn to See Without Images? Procedural Warm-Up for Vision TransformersZachary Shinnick, Liangze Jiang, Hemanth Saratchandran, Damien Teney 等CVPR 2026 · 被引用 9 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 被引用 378 次
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong 等EMNLP 2022 · 被引用 222 次
相关 Paper
- How Useful Is Self-Supervised Pretraining for Visual Tasks?Alejandro Newell, Jia DengCVPR 2020
- Task2Sim: Towards Effective Pre-training and Transfer from Synthetic DataSamarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Richard Chen 等CVPR 2022 · 被引用 28 次
- LIME: Learning Inductive Bias for Primitives of Mathematical ReasoningYuhuai Wu, Markus N. Rabe, Wenda Li, Jimmy Ba 等ICML 2021 · 被引用 66 次
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang 等ICML 2024 · 被引用 61 次
- Task-Robust Pre-Training for Worst-Case Downstream AdaptationJianghui Wang, Yang Chen, Xingyu Xie, Cong Fang 等NeurIPS 2023 · 被引用 3 次
