Pretraining with Artificial Language: Studying Transferable Knowledge in Language Models
Ryokan Ri, Yoshimasa Tsuruoka
摘要
We investigate what kind of structural knowledge learned in neural network encoders is transferable to processing natural language. We design artificial languages with structural properties that mimic natural language, pretrain encoders on the data, and see how much performance the encoder exhibits on downstream tasks in natural language. Our experimental results show that pretraining with an artificial language with a nesting dependency structure provides some knowledge transferable to natural language. A follow-up probing analysis indicates that its success in the transfer is related to the amount of encoded contextual information and what is transferred is the knowledge of position-aware context dependence of language. Our results provide insights into how neural network encoders process human languages and the source of crosslingual transferability of recent multilingual language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng 等NeurIPS 2023 · 被引用 149 次
- A Survey of Deep Learning for Mathematical ReasoningPan Lu, Liang Qiu, Wenhao Yu, Sean Welleck 等ACL 2023 · 被引用 43 次
- Insights into Pre-training via Simpler Synthetic TasksYuhuai Wu, Felix Li, Percy LiangNeurIPS 2022 · 被引用 29 次
- Gloss-Free End-to-End Sign Language TranslationKezhou Lin, Xiaohan Wang, Linchao Zhu, Ke Sun 等ACL 2023 · 被引用 25 次
- Pre-training with Synthetic Data Helps Offline Reinforcement LearningZecheng Wang, Che Wang, Zixuan Dong, Keith W. RossICLR 2024 · 被引用 11 次
它引用的顶会 Paper11
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 被引用 378 次
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau 等EMNLP 2021 · 被引用 177 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
- Identifying Elements Essential for BERT's MultilingualityPhilipp Dufter, Hinrich SchützeEMNLP 2020 · 被引用 44 次
相关 Paper
- Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language ModelsIsabel Papadimitriou, Dan JurafskyEMNLP 2020 · 被引用 40 次
- Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic BiasesMichael Y. Hu, Jackson Petty, Chuan Shi, William Merrill 等ACL 2025
- Low-dimensional Structure in the Space of Language Representations is Reflected in Brain ResponsesRichard J. Antonello, Javier S. Turek, Vy Ai Vo, Alexander HuthNeurIPS 2021 · 被引用 60 次
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 被引用 7 次
- On the Transferability of Pre-trained Language Models: A Study from Artificial DatasetsDavid Cheng-Han Chiang, Hung-Yi LeeAAAI 2022 · 被引用 33 次
