Unifying Image Processing as Visual Prompting Question Answering
Yihao Liu, Xiangyu Chen, Xianzheng Ma, Xintao Wang, Jiantao Zhou, Yu Qiao, Chao Dong
Abstract
Image processing is a fundamental task in computer vision, which aims at enhancing image quality and extracting essential features for subsequent vision applications. Traditionally, task-specific models are developed for individual tasks and designing such models requires distinct expertise. Building upon the success of large language models (LLMs) in natural language processing (NLP), there is a similar trend in computer vision, which focuses on developing large-scale models through pretraining and in-context learning. This paradigm shift reduces the reliance on task-specific models, yielding a powerful unified model to deal with various tasks. However, these advances have predominantly concentrated on high-level vision tasks, with less attention paid to low-level vision tasks. To address this issue, we propose a universal model for general image processing that covers image restoration, image enhancement, image feature extraction tasks, etc. Our proposed framework, named PromptGIP, unifies these diverse image processing tasks within a universal framework. Inspired by NLP question answering (QA) techniques, we employ a visual prompting question answering paradigm. Specifically, we treat the input-output image pair as a structured question-answer sentence, thereby reprogramming the image processing task as a prompting QA problem. PromptGIP can undertake diverse cross-domain tasks using provided visual prompts, eliminating the need for taskspecific finetuning. Capable of handling up to 15 different image processing tasks, PromptGIP represents a versatile and adaptive approach to general image processing. While PromptGIP has demonstrated a certain degree of out-of-domain task generalization, further research is expected to fully explore its more powerful capability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 985f4364-c2e6-4706-b4ad-c26b32a88a2eCited by top-tier papers11
- AWRaCLe: All-Weather Image Restoration Using Visual In-Context LearningSudarshan Rajagopalan, Vishal M. PatelAAAI 2025 · 20 citations
- Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-ExpertsHang Guo, Tao Dai, Yuanchao Bai, Bin Chen et al.NeurIPS 2024 · 14 citations
- Learning A Low-Level Vision Generalist via Visual Task PromptXiangyu Chen, Yihao Liu, Yuandong Pu, Wenlong Zhang et al.ACM MM 2024 · 11 citations
- PairEdit: Learning Semantic Variations for Exemplar-based Image EditingHaoguang Lu, Jiacheng Chen, Zhenguo Yang, Aurele Tohokantche Gnanha et al.NeurIPS 2025 · 8 citations
- UP-Restorer: When Unrolling Meets Prompts for Unified Image RestorationMinghao Liu, Wenhan Yang, Jinyi Luo, Jiaying LiuAAAI 2025 · 7 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- Designing a Practical Degradation Model for Deep Blind Image Super-ResolutionKai Zhang, Jingyun Liang, Luc Van Gool, Radu TimofteICCV 2021 · 898 citations
Related papers
- X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context PromptingZeyi Sun, Ziyang Chu, Pan Zhang, Tong Wu et al.ICCV 2025 · 1 citation
- WeatherGFM: Learning a Weather Generalist Foundation Model via In-context LearningXiangyu Zhao, Zhiwang Zhou, Wenlong Zhang, Yihao Liu et al.ICLR 2025
- Pre-Trained Image Processing TransformerHanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu et al.CVPR 2021
- Towards Unifying Medical Vision-and-Language Pre-training via Soft PromptsZhihong Chen, Shizhe Diao, Benyou Wang, Guanbin Li et al.ICCV 2023 · 50 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
