Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning
Xiang Zhuang, Bin Wu, Jiyu Cui, Kehua Feng, Xiaotong Li, Huabin Xing, Keyan Ding, Qiang Zhang, Huajun Chen
Abstract
Molecular structure elucidation involves deducing a molecule's structure from various types of spectral data, which is crucial in chemical experimental analysis. While large language models (LLMs) have shown remarkable proficiency in analyzing and reasoning through complex tasks, they still encounter substantial challenges in molecular structure elucidation. We identify that these challenges largely stem from LLMs'limited grasp of specialized chemical knowledge. In this work, we introduce a Knowledge-enhanced reasoning framework for Molecular Structure Elucidation (K-MSE), leveraging Monte Carlo Tree Search for test-time scaling as a plugin. Specifically, we construct an external molecular substructure knowledge base to extend the LLMs'coverage of the chemical structure space. Furthermore, we design a specialized molecule-spectrum scorer to act as a reward model for the reasoning process, addressing the issue of inaccurate solution evaluation in LLMs. Experimental results show that our approach significantly boosts performance, particularly gaining more than 20% improvement on both GPT-4o-mini and GPT-4o. Our code is available at https://github.com/HICAI-ZJU/K-MSE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e732dcfb-de10-4c0a-be10-8b87ba681bebCited by top-tier papers1
Ask how each one uses itBuilds on12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
Related papers
- Structural Reasoning Improves Molecular Understanding of LLMYunhui Jang, Jaehyung Kim, Sungsoo AhnACL 2025
- Structured Chemistry Reasoning with Large Language ModelsSiru Ouyang, Zhuosheng Zhang, Bing Yan, Xuan Liu et al.ICML 2024 · 29 citations
- SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-SpectraYunyue Su, Jiahui Chen, Zao Jiang, Zhenyi Zhong et al.ICLR 2026 · 1 citation
- SolverLLM: Leveraging Test-Time Scaling for Optimization Problem via LLM-Guided SearchDong Li, Xujiang Zhao, Linlin Yu, Yanchi Liu et al.NeurIPS 2025 · 15 citations
- ChemAgent: Self-updating Memories in Large Language Models Improves Chemical ReasoningXiangru Tang, Tianyu Hu, Muyang Ye, Yanjun Shao et al.ICLR 2025
