DiagramGPT-Llama3: Enabling Editable, High-Fidelity Diagram Generation with Vision Large Language Models
Yongyuan Chen, Minjie Hong, Boxi Wu, Xicheng Han
Abstract
The automation of diagram generation has gained significant attention in recent years. Previous studies mainly focused on generating diagrams from natural language, but often lacked support for user-friendly editing like drag-and-drop. This paper proposes a novel task: generating editable, high-fidelity diagrams from either text or raster images. It is also among the first to introduce diagram restoration and style transfer in this setting.To tackle these tasks, we constructed the Diagram-mxGraph dataset, covering restoration, text-to-diagram generation, and style transfer. We propose two core innovations: Fine-grained Adaptive Background Suppression (FABS) and Component-Aware Adaptive Loss (CAAL). Leveraging pre-trained Vision Transformers (ViTs) and the Diagram Adapter module, our method aligns diagram features with a Large Language Model (LLM) to output diagrams in editable mxGraph format.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
Related papers
- VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and EditingXiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang et al.ICML 2026 · 3 citations
- Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline DiagramsHaoxiang Hu, Yaokun Li, Zeyuan Huang, Cangjun Gao et al.CVPR 2026
- Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram GenerationZhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li et al.ACM MM 2025 · 3 citations
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li et al.NeurIPS 2025 · 30 citations
- Semantic Document Derendering: SVG Reconstruction via Vision-Language ModelingAdam Hazimeh, Ke Wang, Mark Collier, Gilles Baechler et al.AAAI 2026
