From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing
Jingxuan Wei, Cheng Tan, Qi Chen, Gaowei Wu, Siyuan Li, Zhangyang Gao, Linzhuang Sun, Bihui Yu, Ruifeng Guo
摘要
We introduce the task of text-to-diagram generation, which focuses on creating structured visual representations directly from textual descriptions. Existing approaches in text-to-image and text-to-code generation lack the logical organization and flexibility needed to produce accurate, editable diagrams, often resulting in outputs that are either unstructured or difficult to modify. To address this gap, we introduce DiagramGenBenchmark, a comprehensive evaluation framework encompassing eight distinct diagram categories, including flowcharts, model architecture diagrams, and mind maps. Additionally, we present Dia-gramAgent, an innovative framework with four core modules-Plan Agent, Code Agent, Check Agent, and Diagramto-Code Agent-designed to facilitate both the generation and refinement of complex diagrams. Our extensive experiments, which combine objective metrics with human evaluations, demonstrate that DiagramAgent significantly outperforms existing baseline models in terms of accuracy, structural coherence, and modifiability. This work not only establishes a foundational benchmark for the text-to-diagram generation task but also introduces a powerful toolset to advance research and applications in this emerging area.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- AutoFigure: Generating and Refining Publication-Ready Scientific IllustrationsMinjun Zhu, Zhen Lin, Yixuan Weng, Panzhong Lu 等ICLR 2026 · 被引用 28 次
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary ConstructionsJingxuan Wei, Caijun Jia, Qi Chen, Honghao He 等CVPR 2026 · 被引用 14 次
- Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language DescriptionsShivank Garg, Sankalp Mittal, Manish GuptaICLR 2026 · 被引用 1 次
- ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific InferenceQi Chen, Jingxuan Wei, Zhuoya Yao, Haiguang Wang 等ACM MM 2025 · 被引用 1 次
- GeoLoom: High-quality Geometric Diagram Generation from Textual InputXiaojing Wei, Ting Zhang, Wei He, Jingdong Wang 等ICML 2026
它引用的顶会 Paper6
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 被引用 64 次
- OpenBias: Open-Set Bias Detection in Text-to-Image Generative ModelsMoreno D'Incà, Elia Peruzzo, Massimiliano Mancini, Dejia Xu 等CVPR 2024
相关 Paper
- SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse ParsingTong Zhang, Honglin Lin, Zhou Liu, Chong Chen 等ACL 2026 · 被引用 3 次
- Evaluating LLM-Generated Diagrams as GraphsChumeng Liang, Jiaxuan YouEMNLP 2025
- From Model Diagram to Code: A Benchmark Dataset and Multi-Agent FrameworkMengzhen Wang, Xunbin Huang, Jiayuan Xie, Shukai Ma 等ACM MM 2025 · 被引用 1 次
- Factuality Matters: When Image Generation and Editing Meet Structured VisualsLe Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu 等ICLR 2026 · 被引用 15 次
- VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and EditingXiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang 等ICML 2026 · 被引用 3 次
