VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, Jifeng Dai
Abstract
Large language models (LLMs) have notably accelerated progress towards artificial general intelligence (AGI), with their impressive zero-shot capacity for user-tailored tasks, endowing them with immense potential across a range of applications. However, in the field of computer vision, despite the availability of numerous powerful vision foundation models (VFMs), they are still restricted to tasks in a pre-defined form, struggling to match the open-ended task capabilities of LLMs. In this work, we present an LLM-based framework for vision-centric tasks, termed VisionLLM. This framework provides a unified perspective for vision and language tasks by treating images as a foreign language and aligning vision-centric tasks with language tasks that can be flexibly defined and managed using language instructions. An LLM-based decoder can then make appropriate predictions based on these instructions for open-ended tasks. Extensive experiments show that the proposed VisionLLM can achieve different levels of task customization through language instructions, from fine-grained object-level to coarse-grained task-level customization, all with good results. It's noteworthy that, with a generalist LLM-based framework, our model can achieve over 60% mAP on COCO, on par with detection-specific models. We hope this model can set a new baseline for generalist vision and language models. The demo shall be released based on https://github.com/OpenGVLab/InternGPT. The code shall be released at https://github.com/OpenGVLab/VisionLLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers136
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen et al.ICLR 2026 · 110 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su et al.CVPR 2024
- Advancing Visual Large Language Model for Multi-Granular Versatile PerceptionWentao Xiang, Haoxian Tan, Yujie Zhong, Cong Wei et al.ICCV 2025 · 1 citation
- UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language InterfaceHao Tang, Chen-Wei Xie, Haiyang Wang, Xiaoyi Bao et al.NeurIPS 2025 · 30 citations
- Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual TokenizationYang Jin, Kun Xu, Liwei Chen, Chao Liao et al.ICLR 2024 · 87 citations
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai et al.NeurIPS 2024 · 179 citations
