Protein Language Model Fitness is a Matter of Preference
Cade W. Gordon, Amy X. Lu, Pieter Abbeel
摘要
Leveraging billions of years of evolution, scientists have trained protein language models (pLMs) to understand the sequence and structure space of proteins aiding in the design of more functional proteins. Although they have shown ability to improve efficiency in engineering, it remains unclear if such models capture true biological patterns or artifacts of the training data. We aim to predict the circumstances in which pLMs can successfully perform zero-shot fitness estimation. Our work studies trends observed over hundreds of deep mutational scans across multiple different fitness objectives. We find that the likelihood, or abstractly, implicit preference of a certain protein sequence imbued during pretraining is predictive of fitness prediction capabilities. Both over-preferred and under-preferred wild type sequences harm performance. Using influence functions to causally understand how individual data points increase protein likelihoods, we find that there exists a power law tail due to sequence homology. Lastly, under-performance on low likelihood wild type proteins can be remedied by unsupervised finetuning. These findings that pLM zero-shot fitness estimation can be predicted by the likelihood of the engineered sequence can motivate and improve pLMs' deployment in protein maturation campaigns.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran 等NeurIPS 2025 · 被引用 63 次
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsCharles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan FrazerNeurIPS 2025 · 被引用 12 次
- Shrinking Proteins with DiffusionEthan Baron, Alan Nawzad Amin, Ruben Weitzman, Simon d'Oelsnitz 等ICLR 2026 · 被引用 4 次
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny 等ICLR 2026 · 被引用 1 次
- Steering Protein Language ModelsLong-Kai Huang, Rongyi Zhu, Bing He, Jianhua YaoICML 2025
它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier 等ICML 2021 · 被引用 686 次
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi 等NeurIPS 2022 · 被引用 185 次
相关 Paper
- Non-identifiability and the Blessings of Misspecification in Models of Molecular FitnessEli N. Weinstein, Alan Nawzad Amin, Jonathan Frazer, Debora S. MarksNeurIPS 2022 · 被引用 31 次
- Metalic: Meta-Learning In-Context with Protein Language ModelsJacob Beck, Shikha Surana, Manus McAuliffe, Oliver Bent 等ICLR 2025 · 被引用 3 次
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
- DePLM: Denoising Protein Language Models for Property OptimizationZeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li 等NeurIPS 2024 · 被引用 2 次
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen 等ICML 2024 · 被引用 10 次
