Evolution-Inspired Loss Functions for Protein Representation Learning
Chengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen, Qiang Liu, Daniel Jesus Diaz
摘要
AI-based frameworks for protein engineering use self-supervised learning (SSL) to obtain representations for downstream mutation effect predictions. The most common training objective for these methods is wildtype accuracy: given a sequence or structure where a wildtype residue has been masked, predict the missing amino acid. Wildtype accuracy, however, does not align with the primary goal of protein engineering, which is to suggest a mutation rather than to identify what already appears in nature. Here we present Evolutionary Ranking (EvoRank), a training objective that incorporates evolutionary information derived from multiple sequence alignments (MSAs) to learn more diverse protein representations. Evo-Rank corresponds to ranking amino-acid likelihoods in the probability distribution induced by an MSA. This objective forces models to learn the underlying evolutionary dynamics of a protein. Across a variety of phenotypes and datasets, we demonstrate that EvoRank leads to dramatic improvements in zero-shot performance and can compete with models fine-tuned on experimental data. This is particularly important in protein engineering, where it is expensive to obtain data for fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ambient Proteins - Training Diffusion Models on Noisy StructuresGiannis Daras, Jeffrey Ouyang-Zhang, Krithika Ravishankar, Constantinos Daskalakis 等NeurIPS 2025 · 被引用 1 次
- Distilling Structural Representations into Protein Sequence ModelsJeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl 等ICLR 2025
它引用的顶会 Paper7
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
相关 Paper
- Protein Language Model Fitness is a Matter of PreferenceCade W. Gordon, Amy X. Lu, Pieter AbbeelICLR 2025
- MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingBo Chen, Zhilei Bei, Xingyi Cheng, Pan Li 等NeurIPS 2024 · 被引用 20 次
- RankFlow: Property-aware Transport for Protein OptimizationLu Yu, Wei Xiang, Kang Han, Gaowen Liu 等ICLR 2026
- DePLM: Denoising Protein Language Models for Property OptimizationZeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li 等NeurIPS 2024 · 被引用 2 次
- A Structured Observation Distribution for Generative Biological Sequence Prediction and ForecastingEli N. Weinstein, Debora S. MarksICML 2021 · 被引用 12 次
