VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
Lingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li, Guanhua Chen, Dongdong Zhang, Furu Wei
Abstract
Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a unified framework that seamlessly merges vision and coding language models to empower MLLMs with strong multimodal code generation abilities. Leveraging a task vector-based model merging technique, we integrate a state-of-the-art coding LLM into a strong vision-language backbone, while preserving both visual comprehension and advanced coding skills. To support training and evaluation, we introduce the Multimodal Coding Dataset (MCD), a large-scale and diverse collection of 598k samples, including high-quality HTML code, chart image-code pairs, image-augmented StackOverflow QA, and algorithmic problems. Furthermore, we propose InfiBench-V, a novel and challenging benchmark specifically designed to assess models on visually-rich, real-world programming questions that demand a nuanced understanding of both textual and visual contexts. Extensive experiments show that VisCodex achieves state-of-the-art performance among open-source MLLMs and approaches proprietary models like GPT-4o, highlighting the effectiveness of our model merging strategy and new datasets. Our code and data are available at https://github.com/JackLingjie/VisCodex
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 979f62d5-d1fe-4904-bfa3-fd02873c865dCited by top-tier papers6
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code IntelligenceQiushi Sun, Jingyang Gong, Yang Liu, Qiaosheng Chen et al.ICLR 2026 · 9 citations
- Figma2Code: Automating Multimodal Design to Code in the WildYi Gui, Jiawan Zhang, Yina Wang, Tianran Ma et al.ICLR 2026 · 3 citations
- Hierarchical Process Reward Models are Symbolic Vision LearnersShan Zhang, Aotian Chen, Kai Zou, Jindong Gu et al.CVPR 2026 · 1 citation
- DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram ParsingXingchen Zeng, Zhewei Su, Hengming Zhang, Juyong Jiang et al.ICLR 2026
- Investigating Cross-Modal Skill Injection: Scenarios, Methods, and HyperparametersZhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li et al.ACL 2026
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
- VisCoder2: Building Multi-Language Visualization Coding AgentsYuansheng Ni, Songcheng Cai, Xiangchao Chen, Jiarong Liang et al.ICLR 2026 · 3 citations
- Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data GenerationYue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta et al.ACL 2025
- UnifiedVisual: A Framework for Constructing Unified Vision-Language DatasetsPengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang et al.EMNLP 2025
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch et al.ICML 2024 · 89 citations
