APEX: Approximate-but-exhaustive search for ultra-large combinatorial synthesis libraries
Aryan Pedawi, Jordi Silvestre-Ryan, Bradley Worley, Darren Hsu, Kushal Shah, Elias Stehle, Jingrong Zhang, Izhar Wallach
Abstract
Make-on-demand combinatorial synthesis libraries (CSLs) like Enamine REAL have significantly enabled drug discovery efforts. However, their large size presents a challenge for virtual screening, where the goal is to identify the top compounds in a library according to a computational objective (e.g., optimizing docking score) subject to computational constraints under a limited computational budget. For current library sizes---numbering in the tens of billions of compounds---and scoring functions of interest, a routine virtual screening campaign may be limited to scoring fewer than 0.1% of the available compounds, leaving potentially many high scoring compounds undiscovered. Furthermore, as constraints (and sometimes objectives) change during the course of a virtual screening campaign, existing virtual screening algorithms typically offer little room for amortization. We propose the approximate-but-exhaustive search protocol for CSLs, or APEX. APEX utilizes a neural network surrogate that exploits the structure of CSLs in the prediction of objectives and constraints to make full enumeration on a consumer GPU possible in under a minute, allowing for exact retrieval of approximate top-k sets. To demonstrate APEX's capabilities, we develop a benchmark CSL comprised of more than 10 million compounds, all of which have been annotated with their docking scores on five medically relevant targets along with physicohemical properties measured with RDKit such that, for any objective and set of constraints, the ground truth top-k compounds can be identified and compared against the retrievals from any virtual screening algorithm. We show APEX's consistently strong performance both in retrieval accuracy and runtime compared to alternative methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Projecting Molecules into Synthesizable Chemical SpacesShitong Luo, Wenhao Gao, Zuofan Wu, Jian Peng et al.ICML 2024 · 23 citations
- Parallel Top-K Algorithms on GPU: A Comprehensive Study and New MethodsJingrong Zhang, Akira Naruse, Xipeng Li, Yong WangSC 2023 · 17 citations
- An efficient graph generative model for navigating ultra-large combinatorial synthesis librariesAryan Pedawi, Pawel Gniewek, Chaoyi Chang, Brandon M. Anderson et al.NeurIPS 2022 · 11 citations
Related papers
- Target-Aware Bandit Allocation for Scalable Surrogate Optimization in Chemical SpaceMohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay et al.ICML 2026
- Trillion Ligands per Day: Performance-Portable Virtual Screening via Compound Database Optimization and Multi-Target DockingXiaohui Duan, Cheng Shen, Gaowei Chen, Shanshan Wu et al.SC 2025 · 2 citations
- Learning Compressed Shape-Aware Molecular Representations for Virtual ScreeningRobin Winter, Julian Cremer, Djork-Arné ClevertICML 2026
- Equivariant Scalar Fields for Molecular Docking with Fast Fourier TransformsBowen Jing, Tommi S. Jaakkola, Bonnie BergerICLR 2024 · 3 citations
- DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual ScreeningBowen Gao, Bo Qiang, Haichuan Tan, Yinjun Jia et al.NeurIPS 2023 · 45 citations
