From Model Diagram to Code: A Benchmark Dataset and Multi-Agent Framework
Mengzhen Wang, Xunbin Huang, Jiayuan Xie, Shukai Ma, Jiale Men, Dayong Liang, Yi Cai
Abstract
Model Diagram-to-Code Generation aims to translate model diagrams from research papers into implementation code that reconstructs the model's architecture. This task plays a crucial role in accelerating scientific workflows and enhancing the efficiency of industrial model deployment. While recent studies have explored various Image-to-Code Generation tasks using Multimodal Large Language Models (MLLMs), these efforts have primarily focused on reconstructing the visual appearance depicted in input images, leaving this task largely underexplored. The complex structural elements and implicit relationships in model diagrams present greater challenges for MLLMs, particularly in terms of visual reasoning and semantic interpretation. To support this task, we introduce MDCDataset, a dataset designed to evaluate the ability of MLLMs to generate code from model diagrams. It comprises 1,008 instances spanning 16 research domains, each with a model diagram, structured textual content, and the ground-truth code implementation. Furthermore, to address the inherent challenges of this task, we propose MDCAgent, a collaborative multi-agent framework composed of Parsing, Generation, and Check Agents. These agents work in coordination to analyze, extract, and verify complex elements and implicit relationships within model diagrams, thereby enhancing the visual architecture-aware reasoning capabilities of MLLMs. Our extensive experiments confirm the effectiveness of the framework.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b284c0fa-8822-4f4e-a49a-1946ac9af702Cited by top-tier papers3
- Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal ReasoningDayong Liang, Xiao-Yong Wei, Changmeng ZhengAAAI 2026 · 1 citation
- SRACG: A Code Generation Framework with Selective Retrieval AugmentationMengzhen Wang, Shukai Ma, Songwen Gong, Jiexin Wang et al.AAAI 2026
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge EditingLi Yuan, Qingfei Huang, Bingshan Zhu, Yi Cai et al.AAAI 2026
Related papers
- Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language DescriptionsShivank Garg, Sankalp Mittal, Manish GuptaICLR 2026 · 1 citation
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 86 citations
- mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language ModelAnwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye et al.ACM MM 2024 · 15 citations
- RepLLM: Toward Automatically Reproducing Network Research ResultsYining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu et al.SIGCOMM 2026
- DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram ParsingXingchen Zeng, Zhewei Su, Hengming Zhang, Juyong Jiang et al.ICLR 2026
