S-MolSearch: 3D Semi-supervised Contrastive Learning for Bioactive Molecule Search
Gengmo Zhou, Zhen Wang, Feng Yu, Guolin Ke, Zhewei Wei, Zhifeng Gao
Abstract
Virtual Screening is an essential technique in the early phases of drug discovery, aimed at identifying promising drug candidates from vast molecular libraries. Recently, ligand-based virtual screening has garnered significant attention due to its efficacy in conducting extensive database screenings without relying on specific protein-binding site information. Obtaining binding affinity data for complexes is highly expensive, resulting in a limited amount of available data that covers a relatively small chemical space. Moreover, these datasets contain a significant amount of inconsistent noise. It is challenging to identify an inductive bias that consistently maintains the integrity of molecular activity during data augmentation. To tackle these challenges, we propose S-MolSearch, the first framework to our knowledge, that leverages molecular 3D information and affinity information in semi-supervised contrastive learning for ligand-based virtual screening. Drawing on the principles of inverse optimal transport, S-MolSearch efficiently processes both labeled and unlabeled data, training molecular structural encoders while generating soft labels for the unlabeled data. This design allows S-MolSearch to adaptively utilize unlabeled data within the learning process. Empirically, S-MolSearch demonstrates superior performance on widely-used benchmarks LIT-PCBA and DUD-E. It surpasses both structure-based and ligand-based virtual screening methods for AUROC, BEDROC and EF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccb72fad-1623-444d-a80d-224e563ff4e6Cited by top-tier papers3
- Are High-Degree Representations Really Unnecessary in Equivariant Graph Neural Networks?Jiacheng Cen, Anyi Li, Ning Lin, Yuxiang Ren et al.NeurIPS 2024 · 31 citations
- Universally Invariant Learning in Equivariant GNNsJiacheng Cen, Anyi Li, Ning Lin, Tingyang Xu et al.NeurIPS 2025 · 7 citations
- MIPT: Multilevel Informed Prompt Tuning for Robust Molecular Property PredictionYeyun Chen, Jiangming ShiICML 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
Related papers
- DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual ScreeningBowen Gao, Bo Qiang, Haichuan Tan, Yinjun Jia et al.NeurIPS 2023 · 45 citations
- PharmacoMatch: Efficient 3D Pharmacophore Screening via Neural Subgraph MatchingDaniel Rose, Oliver Wieder, Thomas Seidel, Thierry LangerICLR 2025
- Assay2Mol: Large Language Model-based Drug Design Using BioAssay ContextYifan Deng, Spencer S. Ericksen, Anthony GitterEMNLP 2025 · 1 citation
- S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningBowei He, Bowen Gao, Yankai Chen, Yanyan Lan et al.AAAI 2026 · 1 citation
- Drugging the Undruggable: Benchmarking and Modeling Fragment-Based ScreeningHaichuan Tan, Bowen Gao, Jiaxin Li, Yinjun Jia et al.ICLR 2026
