Self-Supervised Pre-training for Protein Embeddings Using Tertiary Structures
Yuzhi Guo, Jiaxiang Wu, Hehuan Ma, Junzhou Huang
Abstract
The protein tertiary structure largely determines its interaction with other molecules. Despite its importance in various structure-related tasks, fully-supervised data are often timeconsuming and costly to obtain. Existing pre-training models mostly focus on amino-acid sequences or multiple sequence alignments, while the structural information is not yet exploited. In this paper, we propose a self-supervised pre-training model for learning structure embeddings from protein tertiary structures. Native protein structures are perturbed with random noise, and the pre-training model aims at estimating gradients over perturbed 3D structures. Specifically, we adopt SE(3)-invariant features as the model inputs and reconstruct gradients over 3D coordinates with SE(3)equivariance preserved. Such a paradigm avoids the usage of sophisticated SE(3)-equivariant models, and dramatically improves the computational efficiency of pre-training models. We demonstrate the effectiveness of our pre-training model on two downstream tasks, protein structure quality assessment (QA) and protein-protein interaction (PPI) site prediction. Hierarchical structure embeddings are extracted to enhance corresponding prediction models. Extensive experiments indicate that such structure embeddings consistently improve the prediction accuracy for both downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c00de7a-3353-404c-9dc7-8d566f409a18Cited by top-tier papers8
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan et al.ICLR 2024 · 285 citations
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan et al.ICLR 2023 · 40 citations
- Pre-Training Protein Encoder via Siamese Sequence-Structure Diffusion Trajectory PredictionZuobai Zhang, Minghao Xu, Aurélie C. Lozano, Vijil Chenthamarakshan et al.NeurIPS 2023 · 35 citations
- A Hierarchical Training Paradigm for Antibody Structure-sequence Co-designFang Wu, Stan Z. LiNeurIPS 2023 · 27 citations
- Molecular Geometry Pretraining with SE(3)-Invariant Denoising Distance MatchingShengchao Liu, Hongyu Guo, Jian TangICLR 2023 · 17 citations
Builds on3
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- Learning Gradient Fields for Molecular Conformation GenerationChence Shi, Shitong Luo, Minkai Xu, Jian TangICML 2021 · 247 citations
- LieTransformer: Equivariant Self-Attention for Lie GroupsMichael J. Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont et al.ICML 2021 · 132 citations
Related papers
- MAPE-PPI: Towards Effective and Efficient Protein-Protein Interaction Prediction via Microenvironment-Aware Protein EmbeddingLirong Wu, Yijun Tian, Yufei Huang, Siyuan Li et al.ICLR 2024 · 47 citations
- 3D Infomax improves GNNs for Molecular Property PredictionHannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou et al.ICML 2022 · 269 citations
- Pre-Training Protein Bi-level Representation Through Span Mask Strategy On 3D Protein ChainsJiale Zhao, Wanru Zhuang, Jia Song, Yaqi Li et al.ICML 2024 · 9 citations
- Evaluating Representation Learning on the Protein Structure UniverseArian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang et al.ICLR 2024 · 26 citations
- Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein EmbeddingZaifei YANG, Samuel Choi, James KwokICML 2026
