VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
Lingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li, Guanhua Chen, Dongdong Zhang, Furu Wei
摘要
Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a unified framework that seamlessly merges vision and coding language models to empower MLLMs with strong multimodal code generation abilities. Leveraging a task vector-based model merging technique, we integrate a state-of-the-art coding LLM into a strong vision-language backbone, while preserving both visual comprehension and advanced coding skills. To support training and evaluation, we introduce the Multimodal Coding Dataset (MCD), a large-scale and diverse collection of 598k samples, including high-quality HTML code, chart image-code pairs, image-augmented StackOverflow QA, and algorithmic problems. Furthermore, we propose InfiBench-V, a novel and challenging benchmark specifically designed to assess models on visually-rich, real-world programming questions that demand a nuanced understanding of both textual and visual contexts. Extensive experiments show that VisCodex achieves state-of-the-art performance among open-source MLLMs and approaches proprietary models like GPT-4o, highlighting the effectiveness of our model merging strategy and new datasets. Our code and data are available at https://github.com/JackLingjie/VisCodex
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code IntelligenceQiushi Sun, Jingyang Gong, Yang Liu, Qiaosheng Chen 等ICLR 2026 · 被引用 9 次
- Figma2Code: Automating Multimodal Design to Code in the WildYi Gui, Jiawan Zhang, Yina Wang, Tianran Ma 等ICLR 2026 · 被引用 3 次
- Hierarchical Process Reward Models are Symbolic Vision LearnersShan Zhang, Aotian Chen, Kai Zou, Jindong Gu 等CVPR 2026 · 被引用 1 次
- DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram ParsingXingchen Zeng, Zhewei Su, Hengming Zhang, Juyong Jiang 等ICLR 2026
- Investigating Cross-Modal Skill Injection: Scenarios, Methods, and HyperparametersZhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li 等ACL 2026
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi 等WWW 2025 · 被引用 38 次
- VisCoder2: Building Multi-Language Visualization Coding AgentsYuansheng Ni, Songcheng Cai, Xiangchao Chen, Jiarong Liang 等ICLR 2026 · 被引用 3 次
- Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data GenerationYue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta 等ACL 2025
- UnifiedVisual: A Framework for Constructing Unified Vision-Language DatasetsPengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang 等EMNLP 2025
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch 等ICML 2024 · 被引用 89 次
