RetrOrchestrator: A Multi-Step Retrosynthesis Agent Dynamically Orchestrating Single-Step Transition Models
Liao Chang, Luotian Yuan, Yiping Ke, Ying Wei
摘要
Multi-step retrosynthesis planning is a fundamental challenge in organic chemistry, defined by its enormous search space. Existing methods typically formulate it as a Markov Decision Process (MDP) with a fixed choice of transition model (i.e., a single-step retrosynthesis model), and focus on improving how to search through better policies and value functions. However, how the transition space itself is navigated remains largely unexplored. This limitation is particularly urgent given our observation of pronounced skill disparity among single-step prediction models: different models exhibit substantially different performance across molecule states. Motivated by this observation, we introduce RetrOrchestrator, an LLM-powered agent that explicitly accounts for model skill disparity by reframing retrosynthesis planning as a Partially Observable Markov Decision Process (POMDP). By regarding each single-step prediction model as a tool, we further propose a scaffold-aware reinforcement learning algorithm to optimize navigation policy within the transition space. As a result, RetrOrchestrator jointly searches which molecule to expand and which single-step model to apply for the molecule at the current step. Empirically, RetrOrchestrator significantly outperforms static baselines on the Retro*-190 benchmark, achieving a state-of-the-art 94.21% success rate (vs. 9.47% off-the-shelf LLM and 82.63% non-LLM state-dependent router), with 92.49% of solved routes invoking two or more SSRs—evidence that the policy is not collapsing to a single specialist or a static router. The same gain persists on a larger out-of-distribution set (PDB-600), with RetrOrchestrator Pareto-optimal in both wall-clock time and model-query count. Code: https://github.com/ScottLiao920/verl-retro-agent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang 等ICLR 2026 · 被引用 406 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
相关 Paper
- Retro-R1: LLM-based Agentic RetrosynthesisWei Liu, Jiangtao Feng, Hongli Yu, Yuxuan Song 等NeurIPS 2025 · 被引用 8 次
- R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement LearningYiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi 等ACL 2026
- LLM-Augmented Chemical Synthesis and Design Decision ProgramsHaorui Wang, Jeff Guo, Lingkai Kong, Rampi Ramprasad 等ICML 2025
- Retro-Expert: Collaborative Reasoning for Interpretable RetrosynthesisXinyi Li, Sai Wang, Yutian Lin, Yu WuICML 2026 · 被引用 4 次
- Retrosynthetic Planning with Dual Value NetworksGuoqing Liu, Di Xue, Shufang Xie, Yingce Xia 等ICML 2023 · 被引用 24 次
