How much pretraining data do language models need to learn syntax?
Laura Pérez-Mayos, Miguel Ballesteros, Leo Wanner
Abstract
Transformers-based pretrained language models achieve outstanding results in many wellknown NLU benchmarks. However, while pretraining methods are very convenient, they are expensive in terms of time and resources. This calls for a study of the impact of pretraining data size on the knowledge of the models. We explore this impact on the syntactic capabilities of RoBERTa, using models trained on incremental sizes of raw text data. First, we use syntactic structural probes to determine whether models pretrained on more data encode a higher amount of syntactic information. Second, we perform a targeted syntactic evaluation to analyze the impact of pretraining data size on the syntactic generalization performance of the models. Third, we compare the performance of the different models on three downstream applications: part-of-speech tagging, dependency parsing and paraphrase identification. We complement our study with an analysis of the cost-benefit trade-off of training such models. Our experiments show that while models pretrained on more data encode more syntactic knowledge and perform better on downstream applications, they do not always offer a better performance across the different syntactic phenomena and come at a higher financial and environmental cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef3a55c5-9fd5-4602-a7b4-631fe641e575Cited by top-tier papers6
- Analyzing the Mono- and Cross-Lingual Pretraining Dynamics of Multilingual Language ModelsTerra Blevins, Hila Gonen, Luke ZettlemoyerEMNLP 2022 · 13 citations
- Bridging Fairness and Environmental Sustainability in Natural Language ProcessingMarius Hessenthaler, Emma Strubell, Dirk Hovy, Anne LauscherEMNLP 2022 · 8 citations
- SocioProbe: What, When, and Where Language Models Learn about SociodemographicsAnne Lauscher, Federico Bianchi, Samuel R. Bowman, Dirk HovyEMNLP 2022 · 6 citations
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language ModelsYi Zhou, José Camacho-Collados, Danushka BollegalaEMNLP 2023 · 1 citation
Builds on6
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 34 citations
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod et al.ACL 2020 · 21 citations
- Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu et al.EMNLP 2020 · 7 citations
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 7 citations
Related papers
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut et al.ACL 2020 · 168 citations
- Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi et al.EMNLP 2022 · 21 citations
- Syntax-Enhanced Pre-trained ModelZenan Xu, Daya Guo, Duyu Tang, Qinliang Su et al.ACL 2021
- Downstream Datasets Make Surprisingly Good Pretraining CorporaKundan Krishna, Saurabh Garg, Jeffrey P. Bigham, Zachary C. LiptonACL 2023 · 11 citations
