Downstream Datasets Make Surprisingly Good Pretraining Corpora
Kundan Krishna, Saurabh Garg, Jeffrey P. Bigham, Zachary C. Lipton
摘要
For most natural language processing tasks, the dominant practice is to finetune large pretrained transformer models (e.g., BERT) using smaller downstream datasets.Despite the success of this approach, it remains unclear to what extent these gainsare attributable to the massive background corpora employed for pretraining versus to the pretraining objectives themselves. This paper introduces a large-scale study of self-pretraining, where the same (downstream) training data is used for both pretraining and finetuning.In experiments addressing both ELECTRA and RoBERTa models and 10 distinct downstream classification datasets, we observe that self-pretraining rivals standard pretraining on the BookWiki corpus (despite using around 10x–500x less data), outperforming the latter on 7 and 5 datasets, respectively.Surprisingly, these task-specific pretrained models often perform well on other tasks,including the GLUE benchmark. Besides classification tasks, self-pretraining also provides benefits on structured output prediction tasks such as span based question answering and commonsense inference, often providing more than 50% of the performance boosts provided by pretraining on the BookWiki corpus. Our results hint that in many scenarios, performance gains attributable to pretraining are driven primarily by the pretraining objective itself and are not always attributable to the use of external pretraining data in massive amounts.These findings are especially relevant in light of concerns about intellectual property and offensive content in web-scale pretraining data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model PerformanceVishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma 等NeurIPS 2024 · 被引用 101 次
- Neural Data Transformer 2: Multi-context Pretraining for Neural Spiking ActivityJoel Ye, Jennifer L. Collinger, Leila Wehbe, Robert A. GauntNeurIPS 2023 · 被引用 100 次
- Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven PriorsIdo Amos, Jonathan Berant, Ankit GuptaICLR 2024 · 被引用 39 次
- Midtraining Bridges Pretraining and Posttraining DistributionsEmmy Liu, Graham Neubig, Chenyan XiongICML 2026 · 被引用 8 次
- Procedural Pretraining: Warming Up Language Models with Abstract DataLiangze Jiang, Zachary Shinnick, Anton Hengel, Hemanth Saratchandran 等ICML 2026 · 被引用 6 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau 等EMNLP 2021 · 被引用 177 次
- Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyRobert Geirhos, Kristof Meding, Felix A. WichmannNeurIPS 2020 · 被引用 154 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
相关 Paper
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen 等EMNLP 2021 · 被引用 176 次
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 被引用 31 次
