Entropy-Reinforced Planning with Large Language Models for Drug Discovery
Xuefeng Liu, Chih-chan Tien, Peng Ding, Songhao Jiang, Rick L. Stevens
Abstract
The objective of drug discovery is to identify chemical compounds that possess specific pharmaceutical properties toward a binding target. Existing large language models (LLMS) can achieve high token matching scores in terms of likelihood for molecule generation. However, relying solely on LLM decoding often results in the generation of molecules that are either invalid due to a single misused token, or suboptimal due to unbalanced exploration and exploitation as a consequence of the LLM's prior experience. Here we propose ERP, Entropy-Reinforced Planning for Transformer Decoding, which employs an entropy-reinforced planning algorithm to enhance the Transformer decoding process and strike a balance between exploitation and exploration. ERP aims to achieve improvements in multiple properties compared to direct sampling from the Transformer. We evaluated ERP on the SARS-CoV-2 virus (3CLPro) and human cancer cell target protein (RTCB) benchmarks and demonstrated that, in both benchmarks, ERP consistently outperforms the current state-of-the-art algorithm by 1-5 percent, and baselines by 5-10 percent, respectively. Moreover, such improvement is robust across Transformer models trained with different objectives. Finally, to further illustrate the capabilities of ERP, we tested our algorithm on three code generation benchmarks and outperformed the current state-of-the-art approach as well. Our code is publicly available at: https: //github.com/xuefeng-cs/ERP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7e4ee03-05d6-4411-ba25-a183ded303b2Cited by top-tier papers2
- Scaling Small Agents Through Strategy AuctionsLisa Alazraki, Shen, Yoram Bachrach, Akhil MathurICML 2026 · 2 citations
- A Unified Federated Framework for Trajectory Data Preparation via LLMsZhihao Zeng, Ziquan Fang, Wei Shao, Lu Chen et al.ICLR 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Multi-Objective Molecule Generation using Interpretable SubstructuresWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 238 citations
- Learning to Navigate The Synthetically Accessible Chemical Space Using Reinforcement LearningSai Krishna Gottipati, Boris Sattarov, Sufeng Niu, Yashaswi Pathak et al.ICML 2020 · 127 citations
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguagePhilipp Seidl, Andreu Vall, Sepp Hochreiter, Günter KlambauerICML 2023 · 69 citations
Related papers
- Planning with Large Language Models for Code GenerationShun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding et al.ICLR 2023 · 15 citations
- Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent InferenceXuanning Hu, Anchen Li, Qianli Xing, Jinglong Ji et al.WWW 2026
- Efficient Evolutionary Search Over Chemical Space with Large Language ModelsHaorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao et al.ICLR 2025
- De novo Drug Design using Reinforcement Learning with Multiple GPT AgentsXiuyuan Hu, Guoqing Liu, Yang Zhao, Hao ZhangNeurIPS 2023 · 42 citations
- Retrieval-based Controllable Molecule GenerationZichao Wang, Weili Nie, Zhuoran Qiao, Chaowei Xiao et al.ICLR 2023 · 8 citations
