Improving Chemical Understanding of LLMs via SMILES Parsing
Yunhui Jang, Jaehyung Kim, Sungsoo Ahn
Abstract
Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science.A fundamental requirement for these models is the ability to accurately understand molecular structures, commonly encoded in the SMILES representation.However, current LLMs struggle to interpret SMILES, even failing to carry out basic tasks such as counting molecular rings.To address this limitation, we introduce CLEANMOL, a novel framework that formulates SMILES parsing into a suite of clean and deterministic tasks explicitly designed to promote graph-level molecular comprehension.These tasks span from subgraph matching to global graph matching, providing structured supervision aligned with molecular structural properties.We construct a molecular pretraining dataset with adaptive difficulty scoring and pre-train open-source LLMs on these tasks.Our results show that CLEANMOL not only enhances structural comprehension but also achieves the best or competes with the baseline on the Mol-Instructions benchmark. c1ccc(C(F)(F)F)c(N2C(N)=C(C#N)C@HC3=C2CCCC3=O)c1Existence of functional group CC(C)=O?Assemble two fragments?Canonicalize?Length of longest carbon chain?Number of sixmembered rings?(OCC)c([C @H]2C(C#N)=C(N)N (c3ccccc3C(F)(F)F)C3 =C2C(=O)CCC3)c1 CCOc1ccc(OCC)c([C @H]2C(C#N)=C(N)N (c3ccccc3C(F)(F)F)C3 =C2C(=O)CCC3)c1 yes 4 2 Subgraph matching Global graph matching (a) Illustration of SMILES parsing tasks.0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Functional group Ring Chain Canonical Assembly Deepseek-V3-chat GPT-4o CLEANMOL (Ours) (b) Failure of LLMs on SMILES parsing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu et al.ICLR 2024 · 137 citations
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataCanwen Xu, Daya Guo, Nan Duan, Julian J. McAuleyEMNLP 2023 · 112 citations
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke et al.EMNLP 2022 · 112 citations
- BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsQizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu et al.EMNLP 2023 · 40 citations
Related papers
- Improving Large Molecular Language Model via Relation-aware Multimodal CollaborationJinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. KimAAAI 2026
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl et al.ICLR 2026 · 5 citations
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song et al.NeurIPS 2025 · 5 citations
- MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and CorrectionYuyang Wu, Jinhui Ye, Shuhao Zhang, Lu Dai et al.EMNLP 2025 · 1 citation
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 33 citations
