Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models
Francesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang, Alex Xijie Lu
摘要
Large pretrained protein language models (PLMs) have improved protein property and structure prediction from sequences via transfer learning, in which weights and representations from PLMs are repurposed for downstream tasks. Although PLMs have shown great promise, currently there is little understanding of how the features learned by pretraining relate to and are useful for downstream tasks. We perform a systematic analysis of transfer learning using PLMs, conducting 370 experiments across a comprehensive suite of factors including different downstream tasks, architectures, model sizes, model depths, and pretraining time. We observe that while almost all down-stream tasks do benefit from pretrained models compared to naive sequence representations, for the majority of tasks performance does not scale with pretraining, and instead relies on low-level features learned early in pretraining. Our results point to a mismatch between current PLM pretraining paradigms and most applications of these models, indicating a need for better pretraining methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong 等NeurIPS 2024 · 被引用 44 次
- Approximating mutual information of high-dimensional variables using learned representationsGokul Gowri, Xiao-Kang Lun, Allon M. Klein, Peng YinNeurIPS 2024 · 被引用 35 次
- MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingBo Chen, Zhilei Bei, Xingyi Cheng, Pan Li 等NeurIPS 2024 · 被引用 20 次
- FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning ApplicationsKieran Didi, Sarah Alamdari, Alex Lu, Bruce Wittmann 等ICML 2026 · 被引用 5 次
- Reverse Distillation: Consistently Scaling Protein Language Model RepresentationsDarius Catrina, Christian Bepler, Samuel Sledzieski, Rohit SinghICLR 2026 · 被引用 2 次
它引用的顶会 Paper13
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 被引用 654 次
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov 等ICLR 2021 · 被引用 366 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
相关 Paper
- Understanding Transfer Learning of RNA Foundation Models on Downstream TasksYuan Li, Heng Yang, Renzhi Chen, Ke LiICML 2026
- Connecting Pre-trained Language Model and Downstream Task via Properties of RepresentationChenwei Wu, Holden Lee, Rong GeNeurIPS 2023 · 被引用 8 次
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 被引用 147 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
- Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer BehaviorShashank Subramanian, Peter Harrington, Kurt Keutzer, Wahid Bhimji 等NeurIPS 2023 · 被引用 173 次
