Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference
Xuanning Hu, Anchen Li, Qianli Xing, Jinglong Ji, Hao Tuo, Bo Yang
Abstract
Large Language Models (LLMs) possess strong representation and reasoning capabilities, but their application to structure-based drug design (SBDD) is limited by insufficient understanding of protein structures and unpredictable molecular generation. To address these challenges, we propose Exploration-Augmented Latent Inference for LLMs (ELILLM), a framework that reinterprets the LLM generation process as an encoding, latent space exploration, and decoding workflow. ELILLM explicitly explores portions of the design problem beyond the model's current knowledge while using a decoding module to handle familiar regions, generating chemically valid and synthetically reasonable molecules. In our implementation, Bayesian optimization guides the systematic exploration of latent embeddings, and a position-aware surrogate model efficiently predicts binding affinity distributions to inform the search. Knowledge-guided decoding further reduces randomness and effectively imposes chemical validity constraints. We demonstrate ELILLM on the CrossDocked2020 benchmark, showing strong controlled exploration and high binding affinity scores compared with seven baseline methods. These results demonstrate that ELILLM can effectively enhance LLMs' capabilities for SBDD. Our code is available at https://github.com/hxnhxn/ELILLM . CCS Concepts • Computing methodologies → Natural language processing; Search methodologies; • Applied computing → Computational biology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f80bbd86-256f-4ca1-851e-91e238360dcdCited by top-tier papers1
Ask how each one uses itBuilds on18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' et al.ACL 2022 · 332 citations
- A 3D Generative Model for Structure-Based Drug DesignShitong Luo, Jiaqi Guan, Jianzhu Ma, Jian PengNeurIPS 2021 · 302 citations
- Pocket2Mol: Efficient Molecular Sampling Based on 3D Protein PocketsXingang Peng, Shitong Luo, Jiaqi Guan, Qi Xie et al.ICML 2022 · 291 citations
Related papers
- CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMsBowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie et al.NeurIPS 2025 · 4 citations
- Prior-Guided Flow Matching for Target-Aware Molecule Design with Learnable Atom NumberJingyuan Zhou, Hao Qian, Shikui Tu, Lei XuNeurIPS 2025 · 11 citations
- Fragment and Geometry Aware Tokenization of Molecules for Structure-Based Drug Design Using Language ModelsCong Fu, Xiner Li, Blake Olson, Heng Ji et al.ICLR 2025 · 1 citation
- LLM-Augmented Chemical Synthesis and Design Decision ProgramsHaorui Wang, Jeff Guo, Lingkai Kong, Rampi Ramprasad et al.ICML 2025
- RAG-Enhanced Collaborative LLM Agents for Drug DiscoveryNamkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani et al.AAAI 2026 · 21 citations
