From Diagrams to Code: Multilingual Programming with Visual Design
Linzheng Chai, Jian Yang, Shukai Liu, Wei Zhang, Liran WANG, JinKe, Tao Sun, Congnan Liu, Chenchen Zhang, Hualei Zhu, Jiaheng Liu, Xianjie Wu
Abstract
In modern software development, particularly in emerging ``vibe coding'' paradigms, project implementation increasingly begins with visual interactions between users and AI coding assistants, where system architectures are communicated through visual designs before coding. This visual-first approach necessitates AI systems capable of interpreting diagrams across multiple programming languages. However, the development of such systems is severely hindered by the lack of large-scale multimodal training data and evaluation benchmarks. To address these limitations, we present M²C-INSTRUCT, a comprehensive multilingual multimodal instruction-tuning dataset containing over 13.1M samples across 50+ programming languages, designed for visual understanding and diagram interpretation in code generation tasks. We validate our dataset by training M²-CODER, a multilingual multimodal software developer that successfully integrates visual design inputs with textual instructions. We also introduce M²EVAL, a novel multilingual evaluation benchmark for multimodal code generation performance. Experiments show our 7B M²-CODER, performs on par with much larger 70B+ models, confirming the quality and effectiveness of our M²C-INSTRUCT. Together, M²C-INSTRUCT, M²-CODER, and M²EVAL provide essential infrastructure for visual-assisted programming in vibe-coding and visual-interactive development workflows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
Related papers
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
- VisCoder2: Building Multi-Language Visualization Coding AgentsYuansheng Ni, Songcheng Cai, Xiangchao Chen, Jiarong Liang et al.ICLR 2026 · 3 citations
- M2RC-EVAL: Massively Multilingual Repository-level Code Completion EvaluationJiaheng Liu, Ken Deng, Congnan Liu, Jian Yang et al.ACL 2025 · 19 citations
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding ModelsLingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li et al.ICLR 2026 · 15 citations
