When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs
Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftahutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
Abstract
Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evaluation of retrosynthesis performance remains limited. Existing benchmarks and metrics typically rely on published synthetic procedures and Top-K accuracy based on single ground-truth, which does not capture the open-ended nature of real-world synthesis planning. We propose a new benchmarking framework for single-step retrosynthesis that evaluates both general-purpose and chemistry-specialized LLMs using ChemCensor, a novel metric for chemical plausibility. By emphasizing plausibility over exact match, this approach better aligns with human synthesis planning practices. We also introduce CREED, a novel dataset comprising millions of ChemCensor-validated reaction records for LLM training, and use it to train a model that improves over the LLM baselines under this benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce96fcd0-649a-41c0-b1f1-84233529ba23Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu et al.ICLR 2024 · 137 citations
- Training a Scientific Reasoning Model for ChemistrySiddharth Narayanan, James D. Braza, Ryan-Rhys Griffiths, Albert Bou et al.NeurIPS 2025 · 62 citations
- BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsQizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu et al.EMNLP 2023 · 40 citations
- Retrosynthesis Prediction with Local Template RetrievalShufang Xie, Rui Yan, Junliang Guo, Yingce Xia et al.AAAI 2023 · 20 citations
Related papers
- LLM-Augmented Chemical Synthesis and Design Decision ProgramsHaorui Wang, Jeff Guo, Lingkai Kong, Rampi Ramprasad et al.ICML 2025
- Retro-R1: LLM-based Agentic RetrosynthesisWei Liu, Jiangtao Feng, Hongli Yu, Yuxuan Song et al.NeurIPS 2025 · 8 citations
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl et al.ICLR 2026 · 5 citations
- ChemOrch: Empowering LLMs with Chemical Intelligence via Groundbreaking Synthetic InstructionsYue Huang, Zhengzhe Jiang, Xiaonan Luo, Kehan Guo et al.NeurIPS 2025 · 5 citations
- A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to OptimizationZiqing Wang, Kexin Zhang, Zihan Zhao, Yibo Wen et al.ACL 2026 · 10 citations
