MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
Wenrui Zhang, Xinggang Wang, Bin Feng, Wenyu Liu
Abstract
Optical Chemical Structure Recognition (OCSR) plays a pivotal role in modern chemical informatics, enabling the automated conversion of chemical structure images from scientific literature, patents, and educational materials into machine-readable molecular representations. This capability is essential for large-scale chemical data mining, drug discovery pipelines, and Large Language Model (LLM) applications in related domains. However, existing OCSR systems face significant challenges in accurately recognizing stereochemical information due to the subtle visual cues that distinguish stereoisomers, such as wedge and dash bonds, ring conformations, and spatial arrangements. To address these challenges, we propose MolSight, a comprehensive learning framework for OCSR that employs a three-stage training paradigm. In the first stage, we conduct pre-training on large-scale but noisy datasets to endow the model with fundamental perception capabilities for chemical structure images. In the second stage, we perform multi-granularity fine-tuning using datasets with richer supervisory signals, systematically exploring how auxiliary tasks—specifically chemical bond classification and atom localization—contribute to molecular formula recognition. Finally, we employ reinforcement learning for post-training optimization and introduce a novel stereochemical structure dataset. Remarkably, we find that even with MolSight's relatively compact parameter size, the Group Relative Policy Optimization (GRPO) algorithm can further enhance the model's performance on stereomolecular. Through extensive experiments across diverse datasets, our results demonstrate that MolSight achieves state-of-the-art performance in (stereo)chemical optical structure recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong et al.KDD 2022 · 533 citations
- Pre-training Molecular Graph Representation with 3D GeometryShengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby et al.ICLR 2022 · 440 citations
Related papers
- MolParser: End-to-End Visual Recognition of Molecule Structures in the WildXi Fang, Jiankun Wang, Xiaochen Cai, Shangqian Chen et al.ICCV 2025 · 8 citations
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided CaptioningJiahe Song, Chuang Wang, Bowen Jiang, Yinfan Wang et al.CVPR 2026 · 3 citations
- Atom-Level Optical Chemical Structure Recognition with Limited SupervisionMartijn Oldenhof, Edward De Brouwer, Adam Arany, Yves MoreauCVPR 2024 · 2 citations
- Reference-guided Policy Optimization for Molecular Optimization via LLM ReasoningXuan Li, Zhanke Zhou, Zongze Li, Jiangchao Yao et al.ICLR 2026 · 5 citations
- Improving Chemical Understanding of LLMs via SMILES ParsingYunhui Jang, Jaehyung Kim, Sungsoo AhnEMNLP 2025 · 2 citations
