Evolution-Inspired Loss Functions for Protein Representation Learning
Chengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen, Qiang Liu, Daniel Jesus Diaz
Abstract
AI-based frameworks for protein engineering use self-supervised learning (SSL) to obtain representations for downstream mutation effect predictions. The most common training objective for these methods is wildtype accuracy: given a sequence or structure where a wildtype residue has been masked, predict the missing amino acid. Wildtype accuracy, however, does not align with the primary goal of protein engineering, which is to suggest a mutation rather than to identify what already appears in nature. Here we present Evolutionary Ranking (EvoRank), a training objective that incorporates evolutionary information derived from multiple sequence alignments (MSAs) to learn more diverse protein representations. Evo-Rank corresponds to ranking amino-acid likelihoods in the probability distribution induced by an MSA. This objective forces models to learn the underlying evolutionary dynamics of a protein. Across a variety of phenotypes and datasets, we demonstrate that EvoRank leads to dramatic improvements in zero-shot performance and can compete with models fine-tuned on experimental data. This is particularly important in protein engineering, where it is expensive to obtain data for fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Ambient Proteins - Training Diffusion Models on Noisy StructuresGiannis Daras, Jeffrey Ouyang-Zhang, Krithika Ravishankar, Constantinos Daskalakis et al.NeurIPS 2025 · 1 citation
- Distilling Structural Representations into Protein Sequence ModelsJeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl et al.ICLR 2025
Builds on7
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong et al.ICLR 2021 · 357 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
Related papers
- Protein Language Model Fitness is a Matter of PreferenceCade W. Gordon, Amy X. Lu, Pieter AbbeelICLR 2025
- MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingBo Chen, Zhilei Bei, Xingyi Cheng, Pan Li et al.NeurIPS 2024 · 20 citations
- RankFlow: Property-aware Transport for Protein OptimizationLu Yu, Wei Xiang, Kang Han, Gaowen Liu et al.ICLR 2026
- DePLM: Denoising Protein Language Models for Property OptimizationZeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li et al.NeurIPS 2024 · 2 citations
- A Structured Observation Distribution for Generative Biological Sequence Prediction and ForecastingEli N. Weinstein, Debora S. MarksICML 2021 · 12 citations
