Self-supervised Pocket Pretraining via Protein Fragment-Surroundings Alignment
Bowen Gao, Yinjun Jia, Yuanle Mo, Yuyan Ni, Wei-Ying Ma, Zhi-Ming Ma, Yanyan Lan
Abstract
Pocket representations play a vital role in various biomedical applications, such as druggability estimation, ligand affinity prediction, and de novo drug design. While existing geometric features and pretrained representations have demonstrated promising results, they usually treat pockets independent of ligands, neglecting the fundamental interactions between them. However, the limited pocket-ligand complex structures available in the PDB database (less than 100 thousand non-redundant pairs) hampers large-scale pretraining endeavors for interaction modeling. To address this constraint, we propose a novel pocket pretraining approach that leverages knowledge from high-resolution atomic protein structures, assisted by highly effective pretrained small molecule representations. By segmenting protein structures into drug-like fragments and their corresponding pockets, we obtain a reasonable simulation of ligand-receptor interactions, resulting in the generation of over 5 million complexes. Subsequently, the pocket encoder is trained in a contrastive manner to align with the representation of pseudo-ligand furnished by some pretrained small molecule encoders. Our method, named ProFSA, achieves state-of-the-art performance across various tasks, including pocket druggability prediction, pocket matching, and ligand binding affinity prediction. Notably, ProFSA surpasses other pretraining methods by a substantial margin. Moreover, our work opens up a new avenue for mitigating the scarcity of protein-ligand complex data through the utilization of high-quality and diverse protein structure databases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e1a1d90-ea38-4ca4-a091-92b0d1e7993dCited by top-tier papers8
- Accurately Predicting Protein Mutational Effects via a Hierarchical Many-Body Attention NetworkDahao Xu, Jiahua Rao, Mingming Zhu, Jixian Zhang et al.NeurIPS 2025 · 4 citations
- CARD: Coarse-to-fine Autoregressive Modeling with Radix-based Decomposition for Transferable Free Energy EstimationZiyang Yu, Yi He, Wenbing Huang, Wen Yan et al.ICML 2026 · 1 citation
- FIGRDock: Fast Interaction-Guided Regression for Flexible DockingShikun Feng, Bicheng Lin, Yuanhuan Mo, Yuyan Ni et al.NeurIPS 2025
- Redefining the task of Bioactivity PredictionYanwen Huang, Bowen Gao, Yinjun Jia, Hongbo Ma et al.ICLR 2025
- Towards All-Atom Foundation Models for Biomolecular Binding Affinity PredictionLiang Shi, Zuobai Zhang, Huiyu Cai, Santiago Miret et al.ICLR 2026
Builds on12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
- Multi-Scale Representation Learning on ProteinsVignesh Ram Somnath, Charlotte Bunne, Andreas KrauseNeurIPS 2021 · 125 citations
Related papers
- Protein-ligand binding representation learning from fine-grained interactionsShikun Feng, Minghao Li, Yinjun Jia, Wei-Ying Ma et al.ICLR 2024 · 21 citations
- Learning Subpocket Prototypes for Generalizable Structure-based Drug DesignZaixi Zhang, Qi LiuICML 2023 · 42 citations
- Generalized Protein Pocket Generation with Prior-Informed Flow MatchingZaixi Zhang, Marinka Zitnik, Qi LiuNeurIPS 2024 · 10 citations
- S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningBowei He, Bowen Gao, Yankai Chen, Yanyan Lan et al.AAAI 2026 · 1 citation
- Fragment and Geometry Aware Tokenization of Molecules for Structure-Based Drug Design Using Language ModelsCong Fu, Xiner Li, Blake Olson, Heng Ji et al.ICLR 2025 · 1 citation
