LLM-Augmented Chemical Synthesis and Design Decision Programs
Haorui Wang, Jeff Guo, Lingkai Kong, Rampi Ramprasad, Philippe Schwaller, Yuanqi Du, Chao Zhang
Abstract
Retrosynthesis, the process of breaking down a target molecule into simpler precursors through a series of valid reactions, stands at the core of organic chemistry and drug development. Although recent machine learning (ML) research has advanced single-step retrosynthetic modeling and subsequent route searches, these solutions remain restricted by the extensive combinatorial space of possible pathways. Concurrently, large language models (LLMs) have exhibited remarkable chemical knowledge, hinting at their potential to tackle complex decision-making tasks in chemistry. In this work, we explore whether LLMs can successfully navigate the highly constrained, multistep retrosynthesis planning problem. We introduce an efficient scheme for encoding reaction pathways and present a new route-level search strategy, moving beyond the conventional step-bystep reactant prediction. Through comprehensive evaluations, we show that our LLM-augmented approach excels at retrosynthesis planning and extends naturally to the broader challenge of synthesizable molecular design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- A Genetic Algorithm for Navigating Synthesizable Molecular SpacesAlston Lo, Connor W. Coley, Wojciech MatusikICLR 2026 · 6 citations
- When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMsBogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev et al.ICML 2026 · 3 citations
- Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow MatchingChenguang Wang, Zihan Zhou, LEI BAI, Tianshu YuICML 2026 · 1 citation
- R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement LearningYiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi et al.ACL 2026
- RetrOrchestrator: A Multi-Step Retrosynthesis Agent Dynamically Orchestrating Single-Step Transition ModelsLiao Chang, Luotian Yuan, Yiping Ke, Ying WeiICML 2026
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
- MARS: Markov Molecular Sampling for Multi-objective Drug DiscoveryYutong Xie, Chence Shi, Hao Zhou, Yuwei Yang et al.ICLR 2021 · 186 citations
Related papers
- Retro-R1: LLM-based Agentic RetrosynthesisWei Liu, Jiangtao Feng, Hongli Yu, Yuxuan Song et al.NeurIPS 2025 · 8 citations
- Retro-Expert: Collaborative Reasoning for Interpretable RetrosynthesisXinyi Li, Sai Wang, Yutian Lin, Yu WuICML 2026 · 4 citations
- RetroInText: A Multimodal Large Language Model Enhanced Framework for Retrosynthetic Planning via In-Context Representation LearningChenglong Kang, Xiaoyi Liu, Fei GuoICLR 2025
- Towards understanding retrosynthesis by energy-based modelsRuoxi Sun, Hanjun Dai, Li Li, Steven Kearnes et al.NeurIPS 2021 · 45 citations
- A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to OptimizationZiqing Wang, Kexin Zhang, Zihan Zhao, Yibo Wen et al.ACL 2026 · 10 citations
