Venus-MAXWELL: Efficient Learning of Protein-Mutation Stability Landscapes using Protein Language Models
Yuanxi Yu, Fan Jiang, Xinzhu Ma, Liang Zhang, Bozitao Zhong, Wanli Ouyang, Guisheng Fan, Huiqun Yu, Liang Hong, Mingchen Li
Abstract
In-silico prediction of protein mutant stability, measured by the difference in Gibbs free energy change (∆∆G), is fundamental for protein engineering. Current sequence-to-label methods typically employ the two-stage pipeline: (i) encoding mutant sequences using neural networks (e.g., transformers), followed by (ii) the ∆∆G regression from the latent representations. Although these methods have demonstrated promising performance, their dependence on specialized neural network encoders significantly increases the complexity. Additionally, the requirement to individually compute latent representations for each mutant site negatively impacts computational efficiency and poses the risk of overfitting. This work proposes the Venus-MAXWELL framework, which reformulates mutation ∆∆G prediction as a sequence-to-landscape task. In Venus-MAXWELL, mutations of a protein and their corresponding ∆∆G values are organized into a landscape matrix, allowing our framework to learn the ∆∆G landscape of a protein with a single forward and backward pass during training. Besides, to facilitate future works, we also curated a large-scale ∆∆G dataset with strict controls on data leakage and redundancy to ensure robust evaluation. Venus-MAXWELL is compatible with multiple protein language models and enables these models for accurate and efficient ∆∆G prediction. For example, when integrated with the ESM-IF, Venus-MAXWELL achieves higher accuracy than ThermoMPNN with 10× faster in inference speed (despite having 50× more parameters than Ther-moMPNN). The training codes, model weights, and datasets are publicly available at https://github.com/ai4protein/Venus-MAXWELL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 859820ff-04df-4658-853c-e50b10776688Builds on7
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan et al.ICLR 2024 · 285 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
Related papers
- Predicting a Protein's Stability under a Million MutationsJeffrey Ouyang-Zhang, Daniel Jesus Diaz, Adam R. Klivans, Philipp KrähenbühlNeurIPS 2023 · 35 citations
- A Simple yet Effective ΔΔG Predictor is An Unsupervised Antibody Optimizer and ExplainerLirong Wu, Yunfan Liu, Haitao Lin, Yufei Huang et al.ICLR 2025
- VenusX: Unlocking Fine-Grained Functional Understanding of ProteinsYang Tan, Wenrui Gou, Bozitao Zhong, Huiqun Yu et al.ICLR 2026 · 6 citations
- MutaPLM: Protein Language Modeling for Mutation Explanation and EngineeringYizhen Luo, Zikun Nie, Massimo Hong, Suyuan Zhao et al.NeurIPS 2024 · 6 citations
- RankFlow: Property-aware Transport for Protein OptimizationLu Yu, Wei Xiang, Kang Han, Gaowen Liu et al.ICLR 2026
