MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures
Lucas Morin, Valéry Weber, Ahmed Nassar, Gerhard Ingmar Meijer, Luc Van Gool, Yawei Li, Peter W. J. Staar
摘要
The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for chemical structures and Markush structures (chemical structure templates) within patent documents are valuable, e.g., for prior-art search. Advancements have been made in the automatic extraction of chemical structures from text and images, yet the Markush structures remain largely unexplored due to their complex multi-modal nature. In this work, we present MarkushGrapher, a multimodal approach for recognizing Markush structures in documents. Our method jointly encodes text, image, and layout information through a Vision-Text-Layout encoder and an Optical Chemical Structure Recognition vision encoder. These representations are merged and used to autoregressively generate a sequential graph representation of the Markush structure along with a table defining its variable groups. To overcome the lack of real-world training data, we propose a synthetic data generation pipeline that produces a wide range of realistic Markush structures. Additionally, we present M2S, the first annotated benchmark of real-world Markush structures, to advance research on this challenging task. Extensive experiments demonstrate that our approach outperforms state-of-the-art chemistryspecific and general-purpose vision-language models in most evaluation settings. Code, models, and datasets are available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MolParser: End-to-End Visual Recognition of Molecule Structures in the WildXi Fang, Jiankun Wang, Xiaochen Cai, Shangqian Chen 等ICCV 2025 · 被引用 8 次
- MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical StructuresTim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer, Valéry Weber 等CVPR 2026 · 被引用 2 次
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement LearningWenrui Zhang, Xinggang Wang, Bin Feng, Wenyu LiuAAAI 2026 · 被引用 1 次
它引用的顶会 Paper6
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- MolGrapher: Graph-based Visual Recognition of Chemical StructuresLucas Morin, Martin Danelljan, Maria Isabel Agea, Ahmed S. Nassar 等ICCV 2023 · 被引用 32 次
- SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document ConversionAhmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos 等ICCV 2025 · 被引用 9 次
- Atom-Level Optical Chemical Structure Recognition with Limited SupervisionMartijn Oldenhof, Edward De Brouwer, Adam Arany, Yves MoreauCVPR 2024 · 被引用 2 次
- Weakly Supervised Knowledge Transfer with Probabilistic Logical Reasoning for Object DetectionMartijn Oldenhof, Adam Arany, Yves Moreau, Edward De BrouwerICLR 2023
相关 Paper
- MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and GenerationFeiyang Cai, Jiahui Bai, Tao Tang, Guijuan He 等ICLR 2026 · 被引用 10 次
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaJunxian Li, Di Zhang, Xunzhi Wang, Zeying Hao 等AAAI 2025 · 被引用 71 次
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided CaptioningJiahe Song, Chuang Wang, Bowen Jiang, Yinfan Wang 等CVPR 2026 · 被引用 3 次
- Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep ResearchKuicai Dong, Shurui Huang, Fangda Ye, Wei Han 等WWW 2026 · 被引用 4 次
- PatentLMM: Large Multimodal Model for Generating Descriptions for Patent FiguresShreya Shukla, Nakul Sharma, Manish Gupta, Anand MishraAAAI 2025 · 被引用 6 次
