SaProt: Protein Language Modeling with Structure-aware Vocabulary
Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, Fajie Yuan
Abstract
Large-scale protein language models (PLMs), such as the ESM family, have achieved remarkable performance in various downstream tasks related to protein structure and function by undergoing unsupervised training on residue sequences. They have become essential tools for researchers and practitioners in biology. However, a limitation of vanilla PLMs is their lack of explicit consideration for protein structure information, which suggests the potential for further improvement. Motivated by this, we introduce the concept of a "structure-aware vocabulary" that integrates residue tokens with structure tokens. The structure tokens are derived by encoding the 3D structure of proteins using Foldseek. We then propose SaProt, a large-scale general-purpose PLM trained on an extensive dataset comprising approximately 40 million protein sequences and structures. Through extensive evaluation, our SaProt model surpasses well-established and renowned baselines across 10 significant downstream tasks, demonstrating its exceptional capacity and broad applicability. We have made the code 1 , pretrained model, and all relevant materials available at https://github.com/ westlake-repl/SaProt .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fecc0ae8-ec50-4ff4-813b-9754896d9260Cited by top-tier papers44
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Proteus: Exploring Protein Structure Generation for Enhanced Designability and EfficiencyChentong Wang, Yannan Qu, Zhangzhi Peng, Yukai Wang et al.ICML 2024 · 35 citations
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang et al.EMNLP 2024 · 28 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological ViewDeli Chen, Yankai Lin, Wei Li, Peng Li et al.AAAI 2020 · 1,353 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov et al.ICLR 2021 · 366 citations
Related papers
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu et al.ICML 2025
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- Distilling Structural Representations into Protein Sequence ModelsJeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl et al.ICLR 2025
- FoldToken: Learning Protein Language via Vector Quantization and BeyondZhangyang Gao, Cheng Tan, Jue Wang, Yufei Huang et al.AAAI 2025 · 29 citations
- Elucidating the Design Space of Multimodal Protein Language ModelsCheng-Yen Hsieh, Xinyou Wang, Daiheng Zhang, Dongyu Xue et al.ICML 2025
