MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
Tim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer, Valéry Weber, Ahmed Nassar, Peter W. J. Staar
Abstract
Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to recognize molecules represented either in figures or in text independently. However, methods for recognizing chemical structures from multimodal descriptions (Markush structures) lag behind in precision and cannot be used for automatic large-scale processing. In this work, we present MarkushGrapher-2, an end-to-end approach for the multimodal recognition of chemical structures in documents. First, our method employs a dedicated OCR model to extract text from chemical images. Second, the text, image, and layout information are jointly encoded through a Vision-Text-Layout encoder and an Optical Chemical Structure Recognition vision encoder. Finally, the resulting encodings are effectively fused through a two-stage training strategy and used to auto-regressively generate a representation of the Markush structure. To address the lack of training data, we introduce an automatic pipeline for constructing a large-scale dataset of real-world Markush structures. In addition, we present IP5-M, a large manually-annotated benchmark of real-world Markush structures, designed to advance research on this challenging task. Extensive experiments show that our approach substantially outperforms state-of-the-art models in multimodal Markush structure recognition, while maintaining strong performance in molecule structure recognition. Code, models, and datasets are released publicly. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- MolGrapher: Graph-based Visual Recognition of Chemical StructuresLucas Morin, Martin Danelljan, Maria Isabel Agea, Ahmed S. Nassar et al.ICCV 2023 · 32 citations
- SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document ConversionAhmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos et al.ICCV 2025 · 9 citations
- MolParser: End-to-End Visual Recognition of Molecule Structures in the WildXi Fang, Jiankun Wang, Xiaochen Cai, Shangqian Chen et al.ICCV 2025 · 8 citations
- MarkushGrapher: Joint Visual and Textual Recognition of Markush StructuresLucas Morin, Valéry Weber, Ahmed Nassar, Gerhard Ingmar Meijer et al.CVPR 2025
- Unifying Vision, Text, and Layout for Universal Document ProcessingZineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang et al.CVPR 2023
Related papers
- OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General ReasoningJun-Peng Jiang, Shiyin Lu, An-Yang Ji, Yinglun Li et al.ICML 2026
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaJunxian Li, Di Zhang, Xunzhi Wang, Zeying Hao et al.AAAI 2025 · 71 citations
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen et al.CVPR 2026 · 9 citations
- RFL: Simplifying Chemical Structure Recognition with Ring-Free LanguageQikai Chang, Mingjun Chen, Changpeng Pi, Pengfei Hu et al.AAAI 2025 · 1 citation
- Atom-Level Optical Chemical Structure Recognition with Limited SupervisionMartijn Oldenhof, Edward De Brouwer, Adam Arany, Yves MoreauCVPR 2024 · 2 citations
