From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing
Jingxuan Wei, Cheng Tan, Qi Chen, Gaowei Wu, Siyuan Li, Zhangyang Gao, Linzhuang Sun, Bihui Yu, Ruifeng Guo
Abstract
We introduce the task of text-to-diagram generation, which focuses on creating structured visual representations directly from textual descriptions. Existing approaches in text-to-image and text-to-code generation lack the logical organization and flexibility needed to produce accurate, editable diagrams, often resulting in outputs that are either unstructured or difficult to modify. To address this gap, we introduce DiagramGenBenchmark, a comprehensive evaluation framework encompassing eight distinct diagram categories, including flowcharts, model architecture diagrams, and mind maps. Additionally, we present Dia-gramAgent, an innovative framework with four core modules-Plan Agent, Code Agent, Check Agent, and Diagramto-Code Agent-designed to facilitate both the generation and refinement of complex diagrams. Our extensive experiments, which combine objective metrics with human evaluations, demonstrate that DiagramAgent significantly outperforms existing baseline models in terms of accuracy, structural coherence, and modifiability. This work not only establishes a foundational benchmark for the text-to-diagram generation task but also introduces a powerful toolset to advance research and applications in this emerging area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- AutoFigure: Generating and Refining Publication-Ready Scientific IllustrationsMinjun Zhu, Zhen Lin, Yixuan Weng, Panzhong Lu et al.ICLR 2026 · 28 citations
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary ConstructionsJingxuan Wei, Caijun Jia, Qi Chen, Honghao He et al.CVPR 2026 · 14 citations
- Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language DescriptionsShivank Garg, Sankalp Mittal, Manish GuptaICLR 2026 · 1 citation
- ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific InferenceQi Chen, Jingxuan Wei, Zhuoya Yao, Haiguang Wang et al.ACM MM 2025 · 1 citation
- GeoLoom: High-quality Geometric Diagram Generation from Textual InputXiaojing Wei, Ting Zhang, Wei He, Jingdong Wang et al.ICML 2026
Builds on6
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 64 citations
- OpenBias: Open-Set Bias Detection in Text-to-Image Generative ModelsMoreno D'Incà, Elia Peruzzo, Massimiliano Mancini, Dejia Xu et al.CVPR 2024
Related papers
- SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse ParsingTong Zhang, Honglin Lin, Zhou Liu, Chong Chen et al.ACL 2026 · 3 citations
- Evaluating LLM-Generated Diagrams as GraphsChumeng Liang, Jiaxuan YouEMNLP 2025
- From Model Diagram to Code: A Benchmark Dataset and Multi-Agent FrameworkMengzhen Wang, Xunbin Huang, Jiayuan Xie, Shukai Ma et al.ACM MM 2025 · 1 citation
- Factuality Matters: When Image Generation and Editing Meet Structured VisualsLe Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu et al.ICLR 2026 · 15 citations
- VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and EditingXiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang et al.ICML 2026 · 3 citations
