Lune

AAAI2026Top-tier venue

DiagramGPT-Llama3: Enabling Editable, High-Fidelity Diagram Generation with Vision Large Language Models

Yongyuan Chen, Minjie Hong, Boxi Wu, Xicheng Han

2026Year

Abstract

The automation of diagram generation has gained significant attention in recent years. Previous studies mainly focused on generating diagrams from natural language, but often lacked support for user-friendly editing like drag-and-drop. This paper proposes a novel task: generating editable, high-fidelity diagrams from either text or raster images. It is also among the first to introduce diagram restoration and style transfer in this setting.To tackle these tasks, we constructed the Diagram-mxGraph dataset, covering restoration, text-to-diagram generation, and style transfer. We propose two core innovations: Fine-grained Adaptive Background Suppression (FABS) and Component-Aware Adaptive Loss (CAAL). Leveraging pre-trained Vision Transformers (ViTs) and the Diagram Adapter module, our method aligns diagram features with a Large Language Model (LLM) to output diagrams in editable mxGraph format.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines