Pre-Training Protein Bi-level Representation Through Span Mask Strategy On 3D Protein Chains
Jiale Zhao, Wanru Zhuang, Jia Song, Yaqi Li, Shuqi Lu
摘要
In recent years, there has been a surge in the development of 3D structure-based pre-trained protein models, representing a significant advancement over pre-trained protein language models in various downstream tasks. However, most existing structure-based pre-trained models primarily focus on the residue level, i.e., alpha carbon atoms, while ignoring other atoms like side chain atoms. We argue that modeling proteins at both residue and atom levels is important since the side chain atoms can also be crucial for numerous downstream tasks, for example, molecular docking. Nevertheless, we find that naively combining residue and atom information during pre-training typically fails. We identify a key reason is the information leakage caused by the inclusion of atom structure in the input, which renders residue-level pre-training tasks trivial and results in insufficiently expressive residue representations. To address this issue, we introduce a span mask pre-training strategy on 3D protein chains to learn meaningful representations of both residues and atoms. This leads to a simple yet effective approach to learning protein representation suitable for diverse downstream tasks. Extensive experimental results on binding site prediction and function prediction tasks demonstrate our proposed pre-training approach significantly outperforms other methods. Our code will be made public.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-trainingMinghao Xu, Jiaze Song, Keming Wu, Xiangxin Zhou 等ICML 2025
- Multimodal 3D Genome Pre-trainingMinghao Yang, Pengteng Li, Yan Liang, Qianyi Cai 等NeurIPS 2025
- Towards All-Atom Foundation Models for Biomolecular Binding Affinity PredictionLiang Shi, Zuobai Zhang, Huiyu Cai, Santiago Miret 等ICLR 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular DockingGabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay 等ICLR 2023 · 被引用 331 次
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng 等ICLR 2023 · 被引用 254 次
- RotoGrad: Gradient Homogenization in Multitask LearningAdrián Javaloy, Isabel ValeraICLR 2022 · 被引用 114 次
- Self-Supervised Pre-training for Protein Embeddings Using Tertiary StructuresYuzhi Guo, Jiaxiang Wu, Hehuan Ma, Junzhou HuangAAAI 2022 · 被引用 43 次
相关 Paper
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong 等NeurIPS 2024 · 被引用 96 次
- ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular ModelingKangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang 等ICML 2024 · 被引用 17 次
- Multi-level Protein Structure Pre-training via Prompt LearningZeyuan Wang, Qiang Zhang, Shuangwei Hu, Haoran Yu 等ICLR 2023
- Pre-training Sequence, Structure, and Surface Features for Comprehensive Protein Representation LearningYouhan Lee, Hasun Yu, Jaemyung Lee, Jaehoon KimICLR 2024 · 被引用 22 次
- CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding ResiduesLinglin Jing, Sheng Xu, Yifan Wang, Yuzhe Zhou 等AAAI 2024 · 被引用 9 次
