LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval
Yuan Chiang, Elvis Hsieh, Chia-Hong Chou, Janosh Riebesell
Abstract
Materials science research requires multi-step reasoning and precise material informatics retrieval, where minor errors can propagate into significant failures in downstream experiments. Despite their general success, Large Language Models (LLMs) often struggle with hallucinations, handling domain-specific data effectively (e.g., crystal structures), and integrating experimental workflows. To address these challenges, we introduce LLaMP, a hierarchical multi-agent framework designed to emulate the materials science research workflow. The high-level supervisor agent decomposes user requests into sub-tasks and coordinates with specialized assistant agents. These assistant agents handle domain-specific tasks, such as retrieving and processing data from the Materials Project (MP) or conducting simulations as needed. This pipeline facilitates iterative refinement of material property retrieval and enables the simulation of real-world research workflows. To ensure reliability, we propose a novel metric combining uncertainty and confidence estimate to evaluate the selfconsistency of responses from LLaMP and baseline methods. Our experiments demonstrate LLaMP's superior performance in material property retrieval, crystal structure editing, and annealing molecular dynamics simulations using pre-trained interatomic potentials. Unlike prior work focused solely on material property prediction or discovery, LLaMP serves as a foundation for autonomous materials research by combining grounded informatics and enabling iterative experimental processes. Code and live demo are available at https://github.com/chiang-yuan/llamp .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d9851ca-71a7-472b-9aa6-0decf706fc6eBuilds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
Related papers
- AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Material StructuresTaoyuze Lv, Alexander Chen, Fengyu Xie, Chu Wu et al.ICML 2026 · 3 citations
- A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?Agustinus Kristiadi, Felix Strieth-Kalthoff, Marta Skreta, Pascal Poupart et al.ICML 2024 · 55 citations
- RAG-Enhanced Collaborative LLM Agents for Drug DiscoveryNamkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani et al.AAAI 2026 · 21 citations
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition ReasoningCheng Yang, Jiaxuan Lu, Haiyuan Wan, Junchi Yu et al.ICLR 2026 · 13 citations
- MolRAG: Unlocking the Power of Large Language Models for Molecular Property PredictionZiting Xian, Jiawei Gu, Lingbo Li, Shangsong LiangACL 2025 · 7 citations
