When Do You Need Billions of Words of Pretraining Data?
Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman
摘要
NLP is currently dominated by language models like RoBERTa which are pretrained on billions of words. But what exact knowledge or skills do Transformer LMs learn from large-scale pretraining that they cannot learn from less data? To explore this question, we adopt five styles of evaluation: classifier probing, information-theoretic probing, unsupervised relative acceptability judgments, unsupervised language model knowledge probing, and fine-tuning on NLU tasks. We then draw learning curves that track the growth of these different measures of model ability with respect to pretraining data volume using the MiniBERTas, a group of RoBERTa models pretrained on 1M, 10M, 100M and 1B words. We find that these LMs require only about 10M to 100M words to learn to reliably encode most syntactic and semantic features we test. They need a much larger quantity of data in order to acquire enough commonsense knowledge and other skills required to master typical downstream NLU tasks. The results suggest that, while the ability to encode linguistic features is almost certainly necessary for language understanding, it is likely that other, unidentified, forms of knowledge are the major drivers of recent improvements in language understanding among large pretrained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQLHaoyang Li, Jing Zhang, Cuiping Li, Hong ChenAAAI 2023 · 被引用 343 次
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao 等ICLR 2022 · 被引用 237 次
- Should We Be Pre-training? An Argument for End-task Aware Training as an AlternativeLucio M. Dery, Paul Michel, Ameet Talwalkar, Graham NeubigICLR 2022 · 被引用 39 次
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz 等ACL 2022 · 被引用 38 次
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 被引用 31 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont 等ACL 2020 · 被引用 703 次
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
相关 Paper
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 被引用 198 次
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume 等EMNLP 2022 · 被引用 34 次
- Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu 等EMNLP 2020 · 被引用 7 次
- Downstream Datasets Make Surprisingly Good Pretraining CorporaKundan Krishna, Saurabh Garg, Jeffrey P. Bigham, Zachary C. LiptonACL 2023 · 被引用 11 次
- Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-TuningArmen Aghajanyan, Sonal Gupta, Luke ZettlemoyerACL 2021
