Connecting Pre-trained Language Model and Downstream Task via Properties of Representation
Chenwei Wu, Holden Lee, Rong Ge
摘要
Recently, researchers have found that representations learned by large-scale pretrained language models are useful in various downstream tasks. However, there is little theoretical understanding of how pre-training performance is related to downstream task performance. In this paper, we analyze how this performance transfer depends on the properties of the downstream task and the structure of the representations. We consider a log-linear model where a word can be predicted from its context through a network having softmax as its last layer. We show that even if the downstream task is highly structured and depends on a simple function of the hidden representation, there are still cases when a low pre-training loss cannot guarantee good performance on the downstream task. On the other hand, we propose and empirically validate the existence of an "anchor vector" in the representation space, and show that this assumption, together with properties of the downstream task, guarantees performance transfer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Efficient Estimation of Kernel Surrogate Models for Task AttributionZhenshuo Zhang, Minxuan Duan, Hongyang R. ZhangICLR 2026 · 被引用 6 次
- Provable unlearning in topic modeling and downstream tasksStanley Wei, Sadhika Malladi, Sanjeev Arora, Amartya SanyalICLR 2025
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly 等ICLR 2020 · 被引用 559 次
- Predicting What You Already Know Helps: Provable Self-Supervised LearningJason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng ZhuoNeurIPS 2021 · 被引用 219 次
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 被引用 119 次
- A Mathematical Exploration of Why Language Models Help Solve Downstream TasksNikunj Saunshi, Sadhika Malladi, Sanjeev AroraICLR 2021 · 被引用 93 次
相关 Paper
- On Transfer of Adversarial Robustness from Pretraining to Downstream TasksLaura Fee Nern, Harsh Raj, Maurice André Georgi, Yash SharmaNeurIPS 2023 · 被引用 9 次
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang 等ICML 2024 · 被引用 61 次
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 被引用 1 次
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman 等ICLR 2026 · 被引用 8 次
- Scaling Laws for Downstream Task Performance in Machine TranslationBerivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas 等ICLR 2025
