S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual Screening
Bowei He, Bowen Gao, Yankai Chen, Yanyan Lan, Chen Ma, Philip S. Yu, Ya-Qin Zhang, Wei-Ying Ma
摘要
Virtual screening (VS) is an essential task in drug discovery, focusing on the identification of small-molecule ligands that bind to specific protein pockets. Existing deep learning methods, from early regression models to recent contrastive learning approaches, primarily rely on structural data while overlooking protein sequences, which are more accessible and can enhance generalizability. However, directly integrating protein sequences poses challenges due to the redundancy and noise in large-scale protein-ligand datasets. To address these limitations, we propose S 2 Drug, a two-stage framework that explicitly incorporates protein Sequence information and 3D Structure context in protein-ligand contrastive representation learning. In the first stage, we perform protein sequence pretraining on ChemBL using an ESM2-based backbone, combined with a tailored data sampling strategy to reduce redundancy and noise on both protein and ligand sides. In the second stage, we fine-tune on PDBBind by fusing sequence and structure information through a residue-level gating module, while introducing an auxiliary binding site prediction task. This auxiliary task guides the model to accurately localize binding residues within the protein sequence and capture their 3D spatial arrangement, thereby refining protein-ligand matching. Across multiple benchmarks, S 2 Drug consistently improves virtual screening performance and achieves strong results on binding site prediction, demonstrating the value of bridging sequence and structure in contrastive learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular DockingGabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay 等ICLR 2023 · 被引用 331 次
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng 等ICLR 2023 · 被引用 254 次
- Pre-training Sequence, Structure, and Surface Features for Comprehensive Protein Representation LearningYouhan Lee, Hasun Yu, Jaemyung Lee, Jaehoon KimICLR 2024 · 被引用 22 次
- Improving PTM Site Prediction by Coupling of Multi-Granularity Structure and Multi-Scale Sequence RepresentationZhengyi Li, Menglu Li, Lida Zhu, Wen ZhangAAAI 2024 · 被引用 11 次
- Learning Complete Protein Representation by Dynamically Coupling of Sequence and StructureBozhen Hu, Cheng Tan, Jun Xia, Yue Liu 等NeurIPS 2024 · 被引用 6 次
相关 Paper
- AANet: Virtual Screening under Structural Uncertainty via Alignment and AggregationWenyu Zhu, Jianhui Wang, Bowen Gao, Yinjun Jia 等NeurIPS 2025 · 被引用 2 次
- DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual ScreeningBowen Gao, Bo Qiang, Haichuan Tan, Yinjun Jia 等NeurIPS 2023 · 被引用 45 次
- Contrastive Geometric Learning Unlocks Unified Structure- and Ligand-Based Drug DesignLisa Schneckenreiter, Sohvi Luukkonen, Lukas Friedrich, Daniel Kuhn 等ICML 2026
- Protein-ligand binding representation learning from fine-grained interactionsShikun Feng, Minghao Li, Yinjun Jia, Wei-Ying Ma 等ICLR 2024 · 被引用 21 次
- DrugHash: Hashing Based Contrastive Learning for Virtual ScreeningJin Han, Yun Hong, Wu-Jun LiAAAI 2025 · 被引用 4 次
