Pre-Training Protein Encoder via Siamese Sequence-Structure Diffusion Trajectory Prediction
Zuobai Zhang, Minghao Xu, Aurélie C. Lozano, Vijil Chenthamarakshan, Payel Das, Jian Tang
摘要
Self-supervised pre-training methods on proteins have recently gained attention, with most approaches focusing on either protein sequences or structures, neglecting the exploration of their joint distribution, which is crucial for a comprehensive understanding of protein functions by integrating co-evolutionary information and structural characteristics. In this work, inspired by the success of denoising diffusion models in generative tasks, we propose the DiffPreT approach to pre-train a protein encoder by sequence-structure joint diffusion modeling. DiffPreT guides the encoder to recover the native protein sequences and structures from the perturbed ones along the joint diffusion trajectory, which acquires the joint distribution of sequences and structures. Considering the essential protein conformational variations, we enhance DiffPreT by a method called Siamese Diffusion Trajectory Prediction (SiamDiff) to capture the correlation between different conformers of a protein. SiamDiff attains this goal by maximizing the mutual information between representations of diffusion trajectories of structurally-correlated conformers. We study the effectiveness of DiffPreT and SiamDiff on both atom- and residue-level structure-based protein understanding tasks. Experimental results show that the performance of DiffPreT is consistently competitive on all tasks, and SiamDiff achieves new state-of-the-art performance, considering the mean ranks on all tasks. Our implementation is available at https://github.com/DeepGraphLearning/SiamDiff.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Multi-Modal Contrastive Diffusion Model for Therapeutic Peptide GenerationYongkang Wang, Xuan Liu, Feng Huang, Zhankun Xiong 等AAAI 2024 · 被引用 28 次
- Multi-Scale Representation Learning for Protein Fitness PredictionZuobai Zhang, Pascal Notin, Yining Huang, Aurélie C. Lozano 等NeurIPS 2024 · 被引用 17 次
- Pre-Training Protein Bi-level Representation Through Span Mask Strategy On 3D Protein ChainsJiale Zhao, Wanru Zhuang, Jia Song, Yaqi Li 等ICML 2024 · 被引用 9 次
- ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-TrainingLe Zhuo, Zewen Chi, Minghao Xu, Heyan Huang 等ACL 2024
- Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-trainingMinghao Xu, Jiaze Song, Keming Wu, Xiangxin Zhou 等ICML 2025
它引用的顶会 Paper32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
相关 Paper
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue CloudsYeqing Lin, Mohammed AlQuraishiICML 2023 · 被引用 105 次
- Repurposing AlphaFold3-like Protein Folding Models for Antibody Sequence and Structure Co-designNianzu Yang, Songlin Jiang, Jian Ma, Huaijin Wu 等NeurIPS 2025 · 被引用 3 次
- 4D Diffusion for Dynamic Protein Structure Prediction with Reference and Motion GuidanceKaihui Cheng, Ce Liu, Qingkun Su, Jun Wang 等AAAI 2025 · 被引用 7 次
- Learning Diffusion Models with Flexible Representation GuidanceChenyu Wang, Cai Zhou, Sharut Gupta, Johnson Lin 等NeurIPS 2025 · 被引用 12 次
