MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction
Yuyang Wu, Jinhui Ye, Shuhao Zhang, Lu Dai, Yonatan Bisk, Olexandr Isayev
摘要
Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors. This raises important concerns about their robustness and reliability in scientific applications. To support more rigorous evaluation of LLMs in chemical reasoning, we present the MOLERR2FIX benchmark, designed to assess LLMs on error detection and correction in molecular descriptions. Unlike existing benchmarks focused on molecule-totext generation or property prediction, MOL-ERR2FIX emphasizes fine-grained chemical understanding. It tasks LLMs with identifying, localizing, explaining, and revising potential structural and semantic errors in molecular descriptions. Specifically, MOLERR2FIX consists of 1,193 fine-grained annotated error instances. Each instance contains quadruple annotations, i.e,. (error type, span location, the explanation, and the correction). These tasks are intended to reflect the types of reasoning and verification required in real-world chemical communication. Evaluations of current stateof-the-art LLMs reveal notable performance gaps, underscoring the need for more robust chemical reasoning capabilities. MolErr2Fix provides a focused benchmark for evaluating such capabilities and aims to support progress toward more reliable and chemically informed language models. All annotations and accompanying evaluation code are publicly available to facilitate future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu 等ICLR 2024 · 被引用 137 次
- Unifying Molecular and Textual Representations via Multi-task Language ModellingDimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther 等ICML 2023 · 被引用 126 次
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke 等EMNLP 2022 · 被引用 112 次
- Text2Mol: Cross-Modal Molecule Retrieval with Natural Language QueriesCarl Edwards, ChengXiang Zhai, Heng JiEMNLP 2021 · 被引用 79 次
- PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video GenerationQiyao Xue, Xiangyu Yin, Boyuan Yang, Wei GaoCVPR 2025
相关 Paper
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl 等ICLR 2026 · 被引用 5 次
- MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and GenerationFeiyang Cai, Jiahui Bai, Tao Tang, Guijuan He 等ICLR 2026 · 被引用 10 次
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
- MetaBench: A Multi-task Benchmark for Assessing LLMs in MetabolomicsYuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi 等ACL 2026 · 被引用 1 次
