Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagrams
Haoxiang Hu, Yaokun Li, Zeyuan Huang, Cangjun Gao, Qiang He, Qingkun Li, Xiaoming Deng, Cuixia Ma, Yu-Kun Lai, Yong-Jin Liu, Hongan Wang
Abstract
Diagrams are widely used in daily life. However, offline diagrams typically exist in the form of images, lacking structured data representation, which significantly limits their reusability and editability. Current research mainly focuses on supporting basic query tasks for online diagrams and does not meet the semantic understanding and interaction requirements for complex offline diagrams. Although large language models (LLMs) possess powerful reasoning and knowledge integration capabilities, their performance in processing offline diagrams is unsatisfactory due to the inability to accurately understand the structure and content of offline diagrams. To address these issues, we propose Di-agramDiff, a framework consisting of a high-precision diagram reconstruction model and an instance-level diagram element recognition model. The framework converts offline diagrams into standardized data structures, enabling LLMs to transition from unable to understand offline diagrams to intelligent assistants capable of semantic reasoning, logical validation, and efficient diagram editing. To deal with the lack of the dataset, we constructed a dataset containing diagrams, and their corresponding question and answering (Q&A) and editing tasks. Experiments demonstrate that DiagramDiff achieves state-of-the-art performance in diagram reconstruction and recognition tasks, significantly enhancing LLMs' understanding and interaction capabilities with offline diagrams.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a799f2c-1833-4621-aba1-db42ad13b839Builds on8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- InChorus: Designing Consistent Multimodal Interactions for Data Visualization on Tablet DevicesArjun Srinivasan, Bongshin Lee, Nathalie Henry Riche, Steven Mark Drucker et al.CHI 2020 · 78 citations
- A Structured Review of Data Management Technology for Interactive Visualization and AnalysisLeilani Battle, Carlos ScheideggerIEEE VIS 2020 · 40 citations
Related papers
- Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram GenerationZhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li et al.ACM MM 2025 · 3 citations
- CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question AnsweringShaowei Wang, Lingling Zhang, Longji Zhu, Tao Qin et al.CVPR 2024 · 5 citations
- mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language ModelAnwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye et al.ACM MM 2024 · 15 citations
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang et al.SIGIR 2025
- Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training FrameworkYuchen He, Peizhi Ying, Liqi Cheng, Kuilin Peng et al.CHI 2026 · 1 citation
