Connecting Pre-trained Language Model and Downstream Task via Properties of Representation
Chenwei Wu, Holden Lee, Rong Ge
Abstract
Recently, researchers have found that representations learned by large-scale pretrained language models are useful in various downstream tasks. However, there is little theoretical understanding of how pre-training performance is related to downstream task performance. In this paper, we analyze how this performance transfer depends on the properties of the downstream task and the structure of the representations. We consider a log-linear model where a word can be predicted from its context through a network having softmax as its last layer. We show that even if the downstream task is highly structured and depends on a simple function of the hidden representation, there are still cases when a low pre-training loss cannot guarantee good performance on the downstream task. On the other hand, we propose and empirically validate the existence of an "anchor vector" in the representation space, and show that this assumption, together with properties of the downstream task, guarantees performance transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3822dc30-53e8-4839-b9e6-9ae9bc3cac1dCited by top-tier papers2
- Efficient Estimation of Kernel Surrogate Models for Task AttributionZhenshuo Zhang, Minxuan Duan, Hongyang R. ZhangICLR 2026 · 6 citations
- Provable unlearning in topic modeling and downstream tasksStanley Wei, Sadhika Malladi, Sanjeev Arora, Amartya SanyalICLR 2025
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
- Predicting What You Already Know Helps: Provable Self-Supervised LearningJason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng ZhuoNeurIPS 2021 · 219 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- A Mathematical Exploration of Why Language Models Help Solve Downstream TasksNikunj Saunshi, Sadhika Malladi, Sanjeev AroraICLR 2021 · 93 citations
Related papers
- On Transfer of Adversarial Robustness from Pretraining to Downstream TasksLaura Fee Nern, Harsh Raj, Maurice André Georgi, Yash SharmaNeurIPS 2023 · 9 citations
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Scaling Laws for Downstream Task Performance in Machine TranslationBerivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas et al.ICLR 2025
