Self-Supervised Pre-training for Protein Embeddings Using Tertiary Structures
Yuzhi Guo, Jiaxiang Wu, Hehuan Ma, Junzhou Huang
摘要
The protein tertiary structure largely determines its interaction with other molecules. Despite its importance in various structure-related tasks, fully-supervised data are often timeconsuming and costly to obtain. Existing pre-training models mostly focus on amino-acid sequences or multiple sequence alignments, while the structural information is not yet exploited. In this paper, we propose a self-supervised pre-training model for learning structure embeddings from protein tertiary structures. Native protein structures are perturbed with random noise, and the pre-training model aims at estimating gradients over perturbed 3D structures. Specifically, we adopt SE(3)-invariant features as the model inputs and reconstruct gradients over 3D coordinates with SE(3)equivariance preserved. Such a paradigm avoids the usage of sophisticated SE(3)-equivariant models, and dramatically improves the computational efficiency of pre-training models. We demonstrate the effectiveness of our pre-training model on two downstream tasks, protein structure quality assessment (QA) and protein-protein interaction (PPI) site prediction. Hierarchical structure embeddings are extracted to enhance corresponding prediction models. Extensive experiments indicate that such structure embeddings consistently improve the prediction accuracy for both downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan 等ICLR 2024 · 被引用 285 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
- Pre-Training Protein Encoder via Siamese Sequence-Structure Diffusion Trajectory PredictionZuobai Zhang, Minghao Xu, Aurélie C. Lozano, Vijil Chenthamarakshan 等NeurIPS 2023 · 被引用 35 次
- A Hierarchical Training Paradigm for Antibody Structure-sequence Co-designFang Wu, Stan Z. LiNeurIPS 2023 · 被引用 27 次
- Molecular Geometry Pretraining with SE(3)-Invariant Denoising Distance MatchingShengchao Liu, Hongyu Guo, Jian TangICLR 2023 · 被引用 17 次
它引用的顶会 Paper3
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 被引用 1,025 次
- Learning Gradient Fields for Molecular Conformation GenerationChence Shi, Shitong Luo, Minkai Xu, Jian TangICML 2021 · 被引用 247 次
- LieTransformer: Equivariant Self-Attention for Lie GroupsMichael J. Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont 等ICML 2021 · 被引用 132 次
相关 Paper
- MAPE-PPI: Towards Effective and Efficient Protein-Protein Interaction Prediction via Microenvironment-Aware Protein EmbeddingLirong Wu, Yijun Tian, Yufei Huang, Siyuan Li 等ICLR 2024 · 被引用 47 次
- 3D Infomax improves GNNs for Molecular Property PredictionHannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou 等ICML 2022 · 被引用 269 次
- Pre-Training Protein Bi-level Representation Through Span Mask Strategy On 3D Protein ChainsJiale Zhao, Wanru Zhuang, Jia Song, Yaqi Li 等ICML 2024 · 被引用 9 次
- Evaluating Representation Learning on the Protein Structure UniverseArian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang 等ICLR 2024 · 被引用 26 次
- Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein EmbeddingZaifei YANG, Samuel Choi, James KwokICML 2026
