Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagrams
Haoxiang Hu, Yaokun Li, Zeyuan Huang, Cangjun Gao, Qiang He, Qingkun Li, Xiaoming Deng, Cuixia Ma, Yu-Kun Lai, Yong-Jin Liu, Hongan Wang
摘要
Diagrams are widely used in daily life. However, offline diagrams typically exist in the form of images, lacking structured data representation, which significantly limits their reusability and editability. Current research mainly focuses on supporting basic query tasks for online diagrams and does not meet the semantic understanding and interaction requirements for complex offline diagrams. Although large language models (LLMs) possess powerful reasoning and knowledge integration capabilities, their performance in processing offline diagrams is unsatisfactory due to the inability to accurately understand the structure and content of offline diagrams. To address these issues, we propose Di-agramDiff, a framework consisting of a high-precision diagram reconstruction model and an instance-level diagram element recognition model. The framework converts offline diagrams into standardized data structures, enabling LLMs to transition from unable to understand offline diagrams to intelligent assistants capable of semantic reasoning, logical validation, and efficient diagram editing. To deal with the lack of the dataset, we constructed a dataset containing diagrams, and their corresponding question and answering (Q&A) and editing tasks. Experiments demonstrate that DiagramDiff achieves state-of-the-art performance in diagram reconstruction and recognition tasks, significantly enhancing LLMs' understanding and interaction capabilities with offline diagrams.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- InChorus: Designing Consistent Multimodal Interactions for Data Visualization on Tablet DevicesArjun Srinivasan, Bongshin Lee, Nathalie Henry Riche, Steven Mark Drucker 等CHI 2020 · 被引用 78 次
- A Structured Review of Data Management Technology for Interactive Visualization and AnalysisLeilani Battle, Carlos ScheideggerIEEE VIS 2020 · 被引用 40 次
相关 Paper
- Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram GenerationZhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li 等ACM MM 2025 · 被引用 3 次
- CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question AnsweringShaowei Wang, Lingling Zhang, Longji Zhu, Tao Qin 等CVPR 2024 · 被引用 5 次
- mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language ModelAnwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye 等ACM MM 2024 · 被引用 15 次
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang 等SIGIR 2025
- Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training FrameworkYuchen He, Peizhi Ying, Liqi Cheng, Kuilin Peng 等CHI 2026 · 被引用 1 次
