Pretraining with Artificial Language: Studying Transferable Knowledge in Language Models
Ryokan Ri, Yoshimasa Tsuruoka
Abstract
We investigate what kind of structural knowledge learned in neural network encoders is transferable to processing natural language. We design artificial languages with structural properties that mimic natural language, pretrain encoders on the data, and see how much performance the encoder exhibits on downstream tasks in natural language. Our experimental results show that pretraining with an artificial language with a nesting dependency structure provides some knowledge transferable to natural language. A follow-up probing analysis indicates that its success in the transfer is related to the amount of encoded contextual information and what is transferred is the knowledge of position-aware context dependence of language. Our results provide insights into how neural network encoders process human languages and the source of crosslingual transferability of recent multilingual language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7398f68-576e-4a62-8983-cd447bfc118bCited by top-tier papers12
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng et al.NeurIPS 2023 · 149 citations
- A Survey of Deep Learning for Mathematical ReasoningPan Lu, Liang Qiu, Wenhao Yu, Sean Welleck et al.ACL 2023 · 43 citations
- Insights into Pre-training via Simpler Synthetic TasksYuhuai Wu, Felix Li, Percy LiangNeurIPS 2022 · 29 citations
- Gloss-Free End-to-End Sign Language TranslationKezhou Lin, Xiaohan Wang, Linchao Zhu, Ke Sun et al.ACL 2023 · 25 citations
- Pre-training with Synthetic Data Helps Offline Reinforcement LearningZecheng Wang, Che Wang, Zixuan Dong, Keith W. RossICLR 2024 · 11 citations
Builds on11
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- Identifying Elements Essential for BERT's MultilingualityPhilipp Dufter, Hinrich SchützeEMNLP 2020 · 44 citations
Related papers
- Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language ModelsIsabel Papadimitriou, Dan JurafskyEMNLP 2020 · 40 citations
- Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic BiasesMichael Y. Hu, Jackson Petty, Chuan Shi, William Merrill et al.ACL 2025
- Low-dimensional Structure in the Space of Language Representations is Reflected in Brain ResponsesRichard J. Antonello, Javier S. Turek, Vy Ai Vo, Alexander HuthNeurIPS 2021 · 60 citations
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 7 citations
- On the Transferability of Pre-trained Language Models: A Study from Artificial DatasetsDavid Cheng-Han Chiang, Hung-Yi LeeAAAI 2022 · 33 citations
