UViM: A Unified Modeling Approach for Vision with Learned Guiding Codes
Alexander Kolesnikov, André Susano Pinto, Lucas Beyer, Xiaohua Zhai, Jeremiah Harmsen, Neil Houlsby
Abstract
We introduce UViM, a unified approach capable of modeling a wide range of computer vision tasks. In contrast to previous models, UViM has the same functional form for all tasks; it requires no task-specific modifications which require extensive human expertise. The approach involves two components: (I) a base model (feedforward) which is trained to directly predict raw vision outputs, guided by a learned discrete code and (II) a language model (autoregressive) that is trained to generate the guiding code. These components complement each other: the language model is well-suited to modeling structured interdependent data, while the base model is efficient at dealing with high-dimensional outputs. We demonstrate the effectiveness of UViM on three diverse and challenging vision tasks: panoptic segmentation, depth prediction and image colorization, where we achieve competitive and near state-of-the-art results. Our experimental results suggest that UViM is a promising candidate for a unified modeling approach in computer vision. Recently, there have been significant advances in the modeling of complex structured outputs in the context of language generation and (conditional) image generation: autoregressive models [49, 41, 25] , GANs [13], VAE [22], VQVAE [51], diffusion models [45, 18] . However, using such techniques to tackle discriminative problems in a unified way remains under-explored. In this work, we propose a new approach, UViM, capable of modeling many vision tasks, leveraging recent advances in discrete representation learning [51] and language modeling [52] . We show competitive results in three diverse tasks: panoptic segmentation [23], depth prediction [43] and colorization [57] . Crucially, there are no task-specific components required for each task. All of the tasks use the same model and are amenable to transfer learning from standard pre-trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 033a5e20-dcf0-4fab-97c0-b0bce964692eCited by top-tier papers50
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionRui Yang, Lin Song, Yanwei Li, Sijie Zhao et al.NeurIPS 2023 · 340 citations
- A Unified Sequence Interface for Vision TasksTing Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin et al.NeurIPS 2022 · 201 citations
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Palette: Image-to-Image Diffusion ModelsChitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee et al.SIGGRAPH 2022 · 1,638 citations
Related papers
- All in Tokens: Unifying Output Space of Visual Tasks via Soft TokenJia Ning, Chen Li, Zheng Zhang, Chunyu Wang et al.ICCV 2023 · 64 citations
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationYecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang et al.ICLR 2025
- InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision GeneralistsYulu Gan, Sungwoo Park, Alexander Schubert, Anthony Philippakis et al.ICLR 2024 · 34 citations
- UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal TasksJiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi et al.ICLR 2023 · 110 citations
- LAVENDER: Unifying Video-Language Understanding as Masked Language ModelingLinjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin et al.CVPR 2023
