Venus-MAXWELL: Efficient Learning of Protein-Mutation Stability Landscapes using Protein Language Models
Yuanxi Yu, Fan Jiang, Xinzhu Ma, Liang Zhang, Bozitao Zhong, Wanli Ouyang, Guisheng Fan, Huiqun Yu, Liang Hong, Mingchen Li
摘要
In-silico prediction of protein mutant stability, measured by the difference in Gibbs free energy change (∆∆G), is fundamental for protein engineering. Current sequence-to-label methods typically employ the two-stage pipeline: (i) encoding mutant sequences using neural networks (e.g., transformers), followed by (ii) the ∆∆G regression from the latent representations. Although these methods have demonstrated promising performance, their dependence on specialized neural network encoders significantly increases the complexity. Additionally, the requirement to individually compute latent representations for each mutant site negatively impacts computational efficiency and poses the risk of overfitting. This work proposes the Venus-MAXWELL framework, which reformulates mutation ∆∆G prediction as a sequence-to-landscape task. In Venus-MAXWELL, mutations of a protein and their corresponding ∆∆G values are organized into a landscape matrix, allowing our framework to learn the ∆∆G landscape of a protein with a single forward and backward pass during training. Besides, to facilitate future works, we also curated a large-scale ∆∆G dataset with strict controls on data leakage and redundancy to ensure robust evaluation. Venus-MAXWELL is compatible with multiple protein language models and enables these models for accurate and efficient ∆∆G prediction. For example, when integrated with the ESM-IF, Venus-MAXWELL achieves higher accuracy than ThermoMPNN with 10× faster in inference speed (despite having 50× more parameters than Ther-moMPNN). The training codes, model weights, and datasets are publicly available at https://github.com/ai4protein/Venus-MAXWELL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan 等ICLR 2024 · 被引用 285 次
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong 等NeurIPS 2024 · 被引用 96 次
相关 Paper
- Predicting a Protein's Stability under a Million MutationsJeffrey Ouyang-Zhang, Daniel Jesus Diaz, Adam R. Klivans, Philipp KrähenbühlNeurIPS 2023 · 被引用 35 次
- A Simple yet Effective ΔΔG Predictor is An Unsupervised Antibody Optimizer and ExplainerLirong Wu, Yunfan Liu, Haitao Lin, Yufei Huang 等ICLR 2025
- VenusX: Unlocking Fine-Grained Functional Understanding of ProteinsYang Tan, Wenrui Gou, Bozitao Zhong, Huiqun Yu 等ICLR 2026 · 被引用 6 次
- MutaPLM: Protein Language Modeling for Mutation Explanation and EngineeringYizhen Luo, Zikun Nie, Massimo Hong, Suyuan Zhao 等NeurIPS 2024 · 被引用 6 次
- RankFlow: Property-aware Transport for Protein OptimizationLu Yu, Wei Xiang, Kang Han, Gaowen Liu 等ICLR 2026
