Protein Language Model Fitness is a Matter of Preference
Cade W. Gordon, Amy X. Lu, Pieter Abbeel
Abstract
Leveraging billions of years of evolution, scientists have trained protein language models (pLMs) to understand the sequence and structure space of proteins aiding in the design of more functional proteins. Although they have shown ability to improve efficiency in engineering, it remains unclear if such models capture true biological patterns or artifacts of the training data. We aim to predict the circumstances in which pLMs can successfully perform zero-shot fitness estimation. Our work studies trends observed over hundreds of deep mutational scans across multiple different fitness objectives. We find that the likelihood, or abstractly, implicit preference of a certain protein sequence imbued during pretraining is predictive of fitness prediction capabilities. Both over-preferred and under-preferred wild type sequences harm performance. Using influence functions to causally understand how individual data points increase protein likelihoods, we find that there exists a power law tail due to sequence homology. Lastly, under-performance on low likelihood wild type proteins can be remedied by unsupervised finetuning. These findings that pLM zero-shot fitness estimation can be predicted by the likelihood of the engineered sequence can motivate and improve pLMs' deployment in protein maturation campaigns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran et al.NeurIPS 2025 · 63 citations
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsCharles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan FrazerNeurIPS 2025 · 12 citations
- Shrinking Proteins with DiffusionEthan Baron, Alan Nawzad Amin, Ruben Weitzman, Simon d'Oelsnitz et al.ICLR 2026 · 4 citations
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny et al.ICLR 2026 · 1 citation
- Steering Protein Language ModelsLong-Kai Huang, Rongyi Zhu, Bing He, Jianhua YaoICML 2025
Builds on9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier et al.ICML 2021 · 686 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
Related papers
- Non-identifiability and the Blessings of Misspecification in Models of Molecular FitnessEli N. Weinstein, Alan Nawzad Amin, Jonathan Frazer, Debora S. MarksNeurIPS 2022 · 31 citations
- Metalic: Meta-Learning In-Context with Protein Language ModelsJacob Beck, Shikha Surana, Manus McAuliffe, Oliver Bent et al.ICLR 2025 · 3 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- DePLM: Denoising Protein Language Models for Property OptimizationZeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li et al.NeurIPS 2024 · 2 citations
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen et al.ICML 2024 · 10 citations
