Pre-Training Protein Bi-level Representation Through Span Mask Strategy On 3D Protein Chains
Jiale Zhao, Wanru Zhuang, Jia Song, Yaqi Li, Shuqi Lu
Abstract
In recent years, there has been a surge in the development of 3D structure-based pre-trained protein models, representing a significant advancement over pre-trained protein language models in various downstream tasks. However, most existing structure-based pre-trained models primarily focus on the residue level, i.e., alpha carbon atoms, while ignoring other atoms like side chain atoms. We argue that modeling proteins at both residue and atom levels is important since the side chain atoms can also be crucial for numerous downstream tasks, for example, molecular docking. Nevertheless, we find that naively combining residue and atom information during pre-training typically fails. We identify a key reason is the information leakage caused by the inclusion of atom structure in the input, which renders residue-level pre-training tasks trivial and results in insufficiently expressive residue representations. To address this issue, we introduce a span mask pre-training strategy on 3D protein chains to learn meaningful representations of both residues and atoms. This leads to a simple yet effective approach to learning protein representation suitable for diverse downstream tasks. Extensive experimental results on binding site prediction and function prediction tasks demonstrate our proposed pre-training approach significantly outperforms other methods. Our code will be made public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-trainingMinghao Xu, Jiaze Song, Keming Wu, Xiangxin Zhou et al.ICML 2025
- Multimodal 3D Genome Pre-trainingMinghao Yang, Pengteng Li, Yan Liang, Qianyi Cai et al.NeurIPS 2025
- Towards All-Atom Foundation Models for Biomolecular Binding Affinity PredictionLiang Shi, Zuobai Zhang, Huiyu Cai, Santiago Miret et al.ICLR 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular DockingGabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay et al.ICLR 2023 · 331 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
- RotoGrad: Gradient Homogenization in Multitask LearningAdrián Javaloy, Isabel ValeraICLR 2022 · 114 citations
- Self-Supervised Pre-training for Protein Embeddings Using Tertiary StructuresYuzhi Guo, Jiaxiang Wu, Hehuan Ma, Junzhou HuangAAAI 2022 · 43 citations
Related papers
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
- ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular ModelingKangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang et al.ICML 2024 · 17 citations
- Multi-level Protein Structure Pre-training via Prompt LearningZeyuan Wang, Qiang Zhang, Shuangwei Hu, Haoran Yu et al.ICLR 2023
- Pre-training Sequence, Structure, and Surface Features for Comprehensive Protein Representation LearningYouhan Lee, Hasun Yu, Jaemyung Lee, Jaehoon KimICLR 2024 · 22 citations
- CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding ResiduesLinglin Jing, Sheng Xu, Yifan Wang, Yuzhe Zhou et al.AAAI 2024 · 9 citations
