Improving Chemical Understanding of LLMs via SMILES Parsing
Yunhui Jang, Jaehyung Kim, Sungsoo Ahn
摘要
Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science.A fundamental requirement for these models is the ability to accurately understand molecular structures, commonly encoded in the SMILES representation.However, current LLMs struggle to interpret SMILES, even failing to carry out basic tasks such as counting molecular rings.To address this limitation, we introduce CLEANMOL, a novel framework that formulates SMILES parsing into a suite of clean and deterministic tasks explicitly designed to promote graph-level molecular comprehension.These tasks span from subgraph matching to global graph matching, providing structured supervision aligned with molecular structural properties.We construct a molecular pretraining dataset with adaptive difficulty scoring and pre-train open-source LLMs on these tasks.Our results show that CLEANMOL not only enhances structural comprehension but also achieves the best or competes with the baseline on the Mol-Instructions benchmark. c1ccc(C(F)(F)F)c(N2C(N)=C(C#N)C@HC3=C2CCCC3=O)c1Existence of functional group CC(C)=O?Assemble two fragments?Canonicalize?Length of longest carbon chain?Number of sixmembered rings?(OCC)c([C @H]2C(C#N)=C(N)N (c3ccccc3C(F)(F)F)C3 =C2C(=O)CCC3)c1 CCOc1ccc(OCC)c([C @H]2C(C#N)=C(N)N (c3ccccc3C(F)(F)F)C3 =C2C(=O)CCC3)c1 yes 4 2 Subgraph matching Global graph matching (a) Illustration of SMILES parsing tasks.0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Functional group Ring Chain Canonical Assembly Deepseek-V3-chat GPT-4o CLEANMOL (Ours) (b) Failure of LLMs on SMILES parsing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng 等ICLR 2023 · 被引用 254 次
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu 等ICLR 2024 · 被引用 137 次
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataCanwen Xu, Daya Guo, Nan Duan, Julian J. McAuleyEMNLP 2023 · 被引用 112 次
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke 等EMNLP 2022 · 被引用 112 次
- BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsQizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu 等EMNLP 2023 · 被引用 40 次
相关 Paper
- Improving Large Molecular Language Model via Relation-aware Multimodal CollaborationJinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. KimAAAI 2026
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl 等ICLR 2026 · 被引用 5 次
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song 等NeurIPS 2025 · 被引用 5 次
- MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and CorrectionYuyang Wu, Jinhui Ye, Shuhao Zhang, Lu Dai 等EMNLP 2025 · 被引用 1 次
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 被引用 33 次
