A Mathematical Exploration of Why Language Models Help Solve Downstream Tasks
Nikunj Saunshi, Sadhika Malladi, Sanjeev Arora
Abstract
Autoregressive language models, pretrained using large text corpora to do well on next word prediction, have been successful at solving many downstream tasks, even with zero-shot usage. However, there is little theoretical understanding of this success. This paper initiates a mathematical study of this phenomenon for the downstream task of text classification by considering the following questions: (1) What is the intuitive connection between the pretraining task of next word prediction and text classification? (2) How can we mathematically formalize this connection and quantify the benefit of language modeling? For (1), we hypothesize, and verify empirically, that classification tasks of interest can be reformulated as sentence completion tasks, thus making language modeling a meaningful pretraining task. With a mathematical formalization of this hypothesis, we make progress towards (2) and show that language models that are -optimal in cross-entropy (log-perplexity) learn features that can linearly solve such classification tasks with O( √ ) error, thus demonstrating that doing well on language modeling can be beneficial for downstream tasks. We experimentally verify various assumptions and theoretical findings, and also use insights from the analysis to design a new objective function that performs well on some classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0daa039-a8f3-44b9-aff3-dddb7db7a714Cited by top-tier papers33
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- Predicting What You Already Know Helps: Provable Self-Supervised LearningJason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng ZhuoNeurIPS 2021 · 219 citations
- The Learnability of In-Context LearningNoam Wies, Yoav Levine, Amnon ShashuaNeurIPS 2023 · 207 citations
- Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context LearningXinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers et al.NeurIPS 2023 · 206 citations
- Understanding Contrastive Learning Requires Incorporating Inductive BiasesNikunj Saunshi, Jordan T. Ash, Surbhi Goel, Dipendra Misra et al.ICML 2022 · 130 citations
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Predicting What You Already Know Helps: Provable Self-Supervised LearningJason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng ZhuoNeurIPS 2021 · 219 citations
Related papers
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 15 citations
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang et al.AAAI 2024 · 7 citations
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni et al.EMNLP 2020 · 104 citations
- Pre-trained Language Models Can be Fully Zero-Shot LearnersXuandong Zhao, Siqi Ouyang, Zhiguo Yu, Ming Wu et al.ACL 2023 · 22 citations
