ChemEval: A Multi-level and Fine-grained Chemical Capability Evaluation for Large Language Models
Yuqing Huang, Rongyang Zhang, Xuesong He, Xuyang Zhi, Hao Wang, Nuo Chen, Zongbo Liu, Xin Li, Feiyang Xu, Deguang Liu, Huadong Liang, Yi Li
Abstract
The emergence of Large Language Models (LLMs) in chemistry marks a significant advancement in applying artificial intelligence to chemical sciences. While these models show promising potential, their effective application in chemistry demands sophisticated evaluation protocols that address the field's inherent complexities. To bridge this critical gap, we introduce ChemEval, an innovative hierarchical assessment framework specifically designed to evaluate LLMs' capabilities across chemical domains. Our methodology incorporates a distinctive four-tier progression system, spanning from basic chemical concepts to advanced theoretical principles. Sixty-two textual and multimodal tasks are designed to enable researchers to conduct fine-grained analysis of model capabilities and achieve precise evaluation via carefully crafted assessment protocols. The framework integrates carefully curated open-source datasets with expert-validated materials, ensuring both practical relevance and scientific rigor. In our experiments, we evaluated the performance of most main-stream LLMs using both zero-shot and fewshot approaches, with carefully designed examples and prompts. Results indicate that general-purpose LLMs, while proficient in understanding chemical literature and following instructions, struggle with tasks requiring deep chemical expertise. In contrast, chemical LLMs perform better in technical tasks but show limitations in general language processing. These findings highlight both the current limitations and future opportunities for LLMs in chemistry. Our research provides a systematic framework for advancing the application of artificial intelligence in chemical research, potentially facilitating new discoveries in the field. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b88c833-410d-47aa-bbb2-07aa2569e7f6Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific ResearchLiangtai Sun, Yang Han, Zihan Zhao, Da Ma et al.AAAI 2024 · 150 citations
Related papers
- A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to OptimizationZiqing Wang, Kexin Zhang, Zihan Zhao, Yibo Wen et al.ACL 2026 · 10 citations
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaJunxian Li, Di Zhang, Xunzhi Wang, Zeying Hao et al.AAAI 2025 · 71 citations
- ChemOrch: Empowering LLMs with Chemical Intelligence via Groundbreaking Synthetic InstructionsYue Huang, Zhengzhe Jiang, Xiaonan Luo, Kehan Guo et al.NeurIPS 2025 · 5 citations
- SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan et al.ACL 2026
- Curriculum Model Merging: Harmonizing Chemical LLMs for Enhanced Cross-Task GeneralizationBaoyi He, Luotian Yuan, Ying Wei, Fei WuNeurIPS 2025
