Multi-Scale Representation Learning for Protein Fitness Prediction
Zuobai Zhang, Pascal Notin, Yining Huang, Aurélie C. Lozano, Vijil Chenthamarakshan, Debora S. Marks, Payel Das, Jian Tang
摘要
Designing novel functional proteins crucially depends on accurately modeling their fitness landscape. Given the limited availability of functional annotations from wet-lab experiments, previous methods have primarily relied on self-supervised models trained on vast, unlabeled protein sequence or structure datasets. While initial protein representation learning studies solely focused on either sequence or structural features, recent hybrid architectures have sought to merge these modalities to harness their respective strengths. However, these sequence-structure models have so far achieved only incremental improvements when compared to the leading sequence-only approaches, highlighting unresolved challenges effectively leveraging these modalities together. Moreover, the function of certain proteins is highly dependent on the granular aspects of their surface topology, which have been overlooked by prior models. To address these limitations, we introduce the Sequence-Structure-Surface Fitness (S3F) model - a novel multimodal representation learning framework that integrates protein features across several scales. Our approach combines sequence representations from a protein language model with Geometric Vector Perceptron networks encoding protein backbone and detailed surface topology. The proposed method achieves state-of-the-art fitness prediction on the ProteinGym benchmark encompassing 217 substitution deep mutational scanning assays, and provides insights into the determinants of protein function. Our code is at https://github.com/DeepGraphLearning/S3F.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Understanding protein function with a multimodal retrieval-augmented foundation modelTimothy F. Truong Jr., Tristan BeplerNeurIPS 2025 · 被引用 15 次
- Greater than the Sum of Its Parts: Building Substructure into Protein Encoding ModelsRobert Calef, Arthur Liang, Manolis Kellis, Marinka ZitnikICLR 2026 · 被引用 2 次
- Steering Protein Language ModelsLong-Kai Huang, Rongyi Zhu, Bing He, Jianhua YaoICML 2025
- Towards Multiscale Graph-based Protein Learning with Geometric Secondary Structural MotifsShih-Hsin Wang, Yuhao Huang, Taos Transue, Justin M. Baker 等NeurIPS 2025
- (Be Cautious!) Bio-Foundation Models Are Not Yet Robust to Biologically Plausible Perturbations and ML TransformationsJinhao Duan, Ruichen Zhang, Gengwei Zhang, Huaizhi Qu 等ICML 2026
它引用的顶会 Paper15
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier 等ICML 2021 · 被引用 686 次
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend 等ICLR 2021 · 被引用 627 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan 等ICLR 2024 · 被引用 285 次
相关 Paper
- DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein DesignYanting Li, Zikang Wang, Jiyue Jiang, Ziqian Lin 等AAAI 2026
- Fast End-to-End Learning on Protein SurfacesFreyr Sverrisson, Jean Feydy, Bruno E. Correia, Michael M. BronsteinCVPR 2021
- ProtGO: Function-Guided Protein Modeling for Unified Representation LearningBozhen Hu, Cheng Tan, Yongjie Xu, Zhangyang Gao 等NeurIPS 2024 · 被引用 10 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
- Learning Complete Protein Representation by Dynamically Coupling of Sequence and StructureBozhen Hu, Cheng Tan, Jun Xia, Yue Liu 等NeurIPS 2024 · 被引用 6 次
