InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
Yulu Gan, Sungwoo Park, Alexander Schubert, Anthony Philippakis, Ahmed M. Alaa
摘要
Recent advances in generative diffusion models have enabled text-controlled synthesis of realistic and diverse images with impressive quality. Despite these remarkable advances, the application of text-to-image generative models in computer vision for standard visual recognition tasks remains limited. The current de facto approach for these tasks is to design model architectures and loss functions that are tailored to the task at hand. In this paper, we develop a unified language interface for computer vision tasks that abstracts away task-specific design choices and enables task execution by following natural language instructions. Our approach involves casting multiple computer vision tasks as text-to-image generation problems. Here, the text represents an instruction describing the task, and the resulting image is a visually-encoded task output. To train our model, we pool commonly-used computer vision datasets covering a range of tasks, including segmentation, object detection, depth estimation, and classification. We then use a large language model to paraphrase prompt templates that convey the specific tasks to be conducted on each image, and through this process, we create a multi-modal and multi-task training dataset comprising input and output images along with annotated instructions. Following the InstructPix2Pix architecture, we apply instruction-tuning to a text-to-image diffusion model using our constructed dataset, steering its functionality from a generative model to an instruction-guided multi-task vision learner. Experiments demonstrate that our model, dubbed InstructCV, performs competitively compared to other generalist and task-specific vision models. Moreover, it exhibits compelling generalization capabilities to unseen data, categories, and user instructions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai 等NeurIPS 2024 · 被引用 179 次
- PromptFix: You Prompt and We Fix the PhotoYongsheng Yu, Ziyun Zeng, Hang Hua, Jianlong Fu 等NeurIPS 2024 · 被引用 55 次
- GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and EditingRongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang 等NeurIPS 2025 · 被引用 5 次
- X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context PromptingZeyi Sun, Ziyang Chu, Pan Zhang, Tong Wu 等ICCV 2025 · 被引用 1 次
- Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot GeneralizationYang Shen, Xiu-Shen Wei, Yifan Sun, Yuxin Song 等ICML 2025
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Instruct-Imagen: Image Generation with Multi-modal InstructionHexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen 等CVPR 2024
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen 等NeurIPS 2023 · 被引用 128 次
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu 等ICCV 2025 · 被引用 2 次
- InstructDiffusion: A Generalist Modeling Interface for Vision TasksZigang Geng, Binxin Yang, Tiankai Hang, Chen Li 等CVPR 2024
- De-Diffusion Makes Text a Strong Cross-Modal InterfaceChen Wei, Chenxi Liu, Siyuan Qiao, Zhishuai Zhang 等CVPR 2024
