MolParser: End-to-End Visual Recognition of Molecule Structures in the Wild
Xi Fang, Jiankun Wang, Xiaochen Cai, Shangqian Chen, Shuwen Yang, Haoyi Tao, Nan Wang, Lin Yao, Linfeng Zhang, Guolin Ke
Abstract
In recent decades, chemistry publications and patents have increased rapidly. A significant portion of key information is embedded in molecular structure figures, complicating large-scale literature searches and limiting the application of large language models in fields such as biology, chemistry, and pharmaceuticals. The automatic extraction of precise chemical structures is of critical importance. However, the presence of numerous Markush structures in real-world documents, along with variations in molecular image quality, drawing styles, and noise, significantly limits the performance of existing optical chemical structure recognition (OCSR) methods. We present Mol-Parser, a novel end-to-end OCSR method that efficiently and accurately recognizes chemical structures from real-world documents, including difficult Markush structure. We use an extended SMILES encoding rule to annotate our training dataset. Under this rule, we build MolParser-7M, a large-scale OCSR dataset based on our E-SMILES representation. While utilizing a large amount of synthetic data, we employed active learning methods to incorporate substantial in-the-wild data, specifically samples cropped from real patents and scientific literature, into the training process. We trained an end-to-end molecular image captioning model, MolParser, using a curriculum learning approach. MolParser significantly outperforms classical and learningbased methods across most scenarios, with potential for broader downstream applications. The dataset is publicly available in huggingface.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical StructuresTim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer, Valéry Weber et al.CVPR 2026 · 2 citations
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement LearningWenrui Zhang, Xinggang Wang, Bin Feng, Wenyu LiuAAAI 2026 · 1 citation
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
Related papers
- MarkushGrapher: Joint Visual and Textual Recognition of Markush StructuresLucas Morin, Valéry Weber, Ahmed Nassar, Gerhard Ingmar Meijer et al.CVPR 2025
- MolGrapher: Graph-based Visual Recognition of Chemical StructuresLucas Morin, Martin Danelljan, Maria Isabel Agea, Ahmed S. Nassar et al.ICCV 2023 · 32 citations
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided CaptioningJiahe Song, Chuang Wang, Bowen Jiang, Yinfan Wang et al.CVPR 2026 · 3 citations
- Handwritten Chemical Structure Image to Structure-Specific Markup Using Random Conditional Guided DecoderJinshui Hu, Hao Wu, Mingjun Chen, Chenyu Liu et al.ACM MM 2023 · 2 citations
- Atom-Level Optical Chemical Structure Recognition with Limited SupervisionMartijn Oldenhof, Edward De Brouwer, Adam Arany, Yves MoreauCVPR 2024 · 2 citations
