S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual Screening
Bowei He, Bowen Gao, Yankai Chen, Yanyan Lan, Chen Ma, Philip S. Yu, Ya-Qin Zhang, Wei-Ying Ma
Abstract
Virtual screening (VS) is an essential task in drug discovery, focusing on the identification of small-molecule ligands that bind to specific protein pockets. Existing deep learning methods, from early regression models to recent contrastive learning approaches, primarily rely on structural data while overlooking protein sequences, which are more accessible and can enhance generalizability. However, directly integrating protein sequences poses challenges due to the redundancy and noise in large-scale protein-ligand datasets. To address these limitations, we propose S 2 Drug, a two-stage framework that explicitly incorporates protein Sequence information and 3D Structure context in protein-ligand contrastive representation learning. In the first stage, we perform protein sequence pretraining on ChemBL using an ESM2-based backbone, combined with a tailored data sampling strategy to reduce redundancy and noise on both protein and ligand sides. In the second stage, we fine-tune on PDBBind by fusing sequence and structure information through a residue-level gating module, while introducing an auxiliary binding site prediction task. This auxiliary task guides the model to accurately localize binding residues within the protein sequence and capture their 3D spatial arrangement, thereby refining protein-ligand matching. Across multiple benchmarks, S 2 Drug consistently improves virtual screening performance and achieves strong results on binding site prediction, demonstrating the value of bridging sequence and structure in contrastive learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular DockingGabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay et al.ICLR 2023 · 331 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
- Pre-training Sequence, Structure, and Surface Features for Comprehensive Protein Representation LearningYouhan Lee, Hasun Yu, Jaemyung Lee, Jaehoon KimICLR 2024 · 22 citations
- Improving PTM Site Prediction by Coupling of Multi-Granularity Structure and Multi-Scale Sequence RepresentationZhengyi Li, Menglu Li, Lida Zhu, Wen ZhangAAAI 2024 · 11 citations
- Learning Complete Protein Representation by Dynamically Coupling of Sequence and StructureBozhen Hu, Cheng Tan, Jun Xia, Yue Liu et al.NeurIPS 2024 · 6 citations
Related papers
- AANet: Virtual Screening under Structural Uncertainty via Alignment and AggregationWenyu Zhu, Jianhui Wang, Bowen Gao, Yinjun Jia et al.NeurIPS 2025 · 2 citations
- DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual ScreeningBowen Gao, Bo Qiang, Haichuan Tan, Yinjun Jia et al.NeurIPS 2023 · 45 citations
- Contrastive Geometric Learning Unlocks Unified Structure- and Ligand-Based Drug DesignLisa Schneckenreiter, Sohvi Luukkonen, Lukas Friedrich, Daniel Kuhn et al.ICML 2026
- Protein-ligand binding representation learning from fine-grained interactionsShikun Feng, Minghao Li, Yinjun Jia, Wei-Ying Ma et al.ICLR 2024 · 21 citations
- DrugHash: Hashing Based Contrastive Learning for Virtual ScreeningJin Han, Yun Hong, Wu-Jun LiAAAI 2025 · 4 citations
