Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models
Francesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang, Alex Xijie Lu
Abstract
Large pretrained protein language models (PLMs) have improved protein property and structure prediction from sequences via transfer learning, in which weights and representations from PLMs are repurposed for downstream tasks. Although PLMs have shown great promise, currently there is little understanding of how the features learned by pretraining relate to and are useful for downstream tasks. We perform a systematic analysis of transfer learning using PLMs, conducting 370 experiments across a comprehensive suite of factors including different downstream tasks, architectures, model sizes, model depths, and pretraining time. We observe that while almost all down-stream tasks do benefit from pretrained models compared to naive sequence representations, for the majority of tasks performance does not scale with pretraining, and instead relies on low-level features learned early in pretraining. Our results point to a mismatch between current PLM pretraining paradigms and most applications of these models, indicating a need for better pretraining methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong et al.NeurIPS 2024 · 44 citations
- Approximating mutual information of high-dimensional variables using learned representationsGokul Gowri, Xiao-Kang Lun, Allon M. Klein, Peng YinNeurIPS 2024 · 35 citations
- MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingBo Chen, Zhilei Bei, Xingyi Cheng, Pan Li et al.NeurIPS 2024 · 20 citations
- FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning ApplicationsKieran Didi, Sarah Alamdari, Alex Lu, Bruce Wittmann et al.ICML 2026 · 5 citations
- Reverse Distillation: Consistently Scaling Protein Language Model RepresentationsDarius Catrina, Christian Bepler, Samuel Sledzieski, Rohit SinghICLR 2026 · 2 citations
Builds on13
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov et al.ICLR 2021 · 366 citations
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong et al.ICLR 2021 · 357 citations
Related papers
- Understanding Transfer Learning of RNA Foundation Models on Downstream TasksYuan Li, Heng Yang, Renzhi Chen, Ke LiICML 2026
- Connecting Pre-trained Language Model and Downstream Task via Properties of RepresentationChenwei Wu, Holden Lee, Rong GeNeurIPS 2023 · 8 citations
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 147 citations
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan et al.ICLR 2023 · 40 citations
- Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer BehaviorShashank Subramanian, Peter Harrington, Kurt Keutzer, Wahid Bhimji et al.NeurIPS 2023 · 173 citations
