Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt Tuning
Colin Wei, Sang Michael Xie, Tengyu Ma
摘要
Pretrained language models have achieved state-of-the-art performance when adapted to a downstream NLP task. However, theoretical analysis of these models is scarce and challenging since the pretraining and downstream tasks can be very different. We propose an analysis framework that links the pretraining and downstream tasks with an underlying latent variable generative model of text -- the downstream classifier must recover a function of the posterior distribution over the latent variables. We analyze head tuning (learning a classifier on top of the frozen pretrained model) and prompt tuning in this setting. The generative model in our analysis is either a Hidden Markov Model (HMM) or an HMM augmented with a latent memory component, motivated by long-term dependencies in natural language. We show that 1) under certain non-degeneracy conditions on the HMM, simple classification heads can solve the downstream task, 2) prompt tuning obtains downstream guarantees with weaker non-degeneracy conditions, and 3) our recovery guarantees for the memory-augmented HMM are stronger than for the vanilla HMM because task-relevant information is easier to recover from the long-term memory. Experiments on synthetically generated data from HMMs back our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- The Learnability of In-Context LearningNoam Wies, Yoav Levine, Amnon ShashuaNeurIPS 2023 · 被引用 207 次
- Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context LearningXinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers 等NeurIPS 2023 · 被引用 206 次
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 被引用 190 次
- Universal Prompt Tuning for Graph Neural NetworksTaoran Fang, Yunchao Zhang, Yang Yang, Chunping Wang 等NeurIPS 2023 · 被引用 166 次
它引用的顶会 Paper14
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 被引用 425 次
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 被引用 261 次
- Predicting What You Already Know Helps: Provable Self-Supervised LearningJason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng ZhuoNeurIPS 2021 · 被引用 219 次
相关 Paper
- On the Role of Attention in Prompt-tuningSamet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, Christos ThrampoulidisICML 2023 · 被引用 67 次
- PPT: Pre-trained Prompt Tuning for Few-shot LearningYuxian Gu, Xu Han, Zhiyuan Liu, Minlie HuangACL 2022
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik 等ICLR 2024 · 被引用 45 次
- Universality and Limitations of Prompt TuningYihan Wang, Jatin Chauhan, Wei Wang, Cho-Jui HsiehNeurIPS 2023 · 被引用 48 次
- HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text ClassificationZihan Wang, Peiyi Wang, Tianyu Liu, Binghuai Lin 等EMNLP 2022 · 被引用 42 次
