SaProt: Protein Language Modeling with Structure-aware Vocabulary
Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, Fajie Yuan
摘要
Large-scale protein language models (PLMs), such as the ESM family, have achieved remarkable performance in various downstream tasks related to protein structure and function by undergoing unsupervised training on residue sequences. They have become essential tools for researchers and practitioners in biology. However, a limitation of vanilla PLMs is their lack of explicit consideration for protein structure information, which suggests the potential for further improvement. Motivated by this, we introduce the concept of a "structure-aware vocabulary" that integrates residue tokens with structure tokens. The structure tokens are derived by encoding the 3D structure of proteins using Foldseek. We then propose SaProt, a large-scale general-purpose PLM trained on an extensive dataset comprising approximately 40 million protein sequences and structures. Through extensive evaluation, our SaProt model surpasses well-established and renowned baselines across 10 significant downstream tasks, demonstrating its exceptional capacity and broad applicability. We have made the code 1 , pretrained model, and all relevant materials available at https://github.com/ westlake-repl/SaProt .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper44
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong 等NeurIPS 2024 · 被引用 96 次
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang 等ICML 2024 · 被引用 61 次
- Proteus: Exploring Protein Structure Generation for Enhanced Designability and EfficiencyChentong Wang, Yannan Qu, Zhangzhi Peng, Yukai Wang 等ICML 2024 · 被引用 35 次
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological ViewDeli Chen, Yankai Lin, Wei Li, Peng Li 等AAAI 2020 · 被引用 1,353 次
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov 等ICLR 2021 · 被引用 366 次
相关 Paper
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu 等ICML 2025
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- Distilling Structural Representations into Protein Sequence ModelsJeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl 等ICLR 2025
- FoldToken: Learning Protein Language via Vector Quantization and BeyondZhangyang Gao, Cheng Tan, Jue Wang, Yufei Huang 等AAAI 2025 · 被引用 29 次
- Elucidating the Design Space of Multimodal Protein Language ModelsCheng-Yen Hsieh, Xinyou Wang, Daiheng Zhang, Dongyu Xue 等ICML 2025
