Learning A Low-Level Vision Generalist via Visual Task Prompt
Xiangyu Chen, Yihao Liu, Yuandong Pu, Wenlong Zhang, Jiantao Zhou, Yu Qiao, Chao Dong
摘要
Building a unified model for general low-level vision tasks holds significant research and practical value. Current methods encounter several critical issues. Multi-task restoration approaches can address multiple degradation-to-clean restoration tasks, while their applicability to tasks with different target domains (e.g., image stylization) is limited. Methods like PromptGIP can handle multiple input-target domains but rely on the Masked Autoencoder (MAE) paradigm. Consequently, they are tied to the ViT architecture, resulting in suboptimal image reconstruction quality. In addition, these methods are sensitive to prompt image content and often struggle with low-frequency information processing. In this paper, we propose a Visual task Prompt-based Image Processing (VPIP) framework to overcome these challenges. VPIP employs visual task prompts to manage tasks with different input-target domains and allows flexible selection of backbone network suitable for general tasks. Besides, a new prompt cross-attention is introduced to facilitate interaction between the input and prompt information. Based on the VPIP framework, we train a low-level vision generalist model, namely GenLV, on 30 diverse tasks. Experimental results show that GenLV can successfully address a variety of low-level tasks, significantly outperforming existing methods both quantitatively and qualitatively. Codes are available at https://github.com/chxy95/GenLV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Factuality Matters: When Image Generation and Editing Meet Structured VisualsLe Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu 等ICLR 2026 · 被引用 15 次
- FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation ModelXiang Chen, Jinshan Pan, Jiangxin Dong, Jian Yang 等CVPR 2026 · 被引用 10 次
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image RestorationJingren Liu, Shuning Xu, Qirui Yang, Yun Wang 等CVPR 2026 · 被引用 4 次
- LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior SamplingHuaqiu Li, Yong Wang, Tongwen Huang, Hailang Huang 等ICCV 2025 · 被引用 4 次
- An Intelligent Agentic System for Complex Image Restoration ProblemsKaiwen Zhu, Jinjin Gu, Zhiyuan You, Yu Qiao 等ICLR 2025
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
相关 Paper
- Unifying Image Processing as Visual Prompting Question AnsweringYihao Liu, Xiangyu Chen, Xianzheng Ma, Xintao Wang 等ICML 2024 · 被引用 37 次
- Controlling Vision-Language Models for Multi-Task Image RestorationZiwei Luo, Fredrik K. Gustafsson, Zheng Zhao, Jens Sjölund 等ICLR 2024 · 被引用 111 次
- X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context PromptingZeyi Sun, Ziyang Chu, Pan Zhang, Tong Wu 等ICCV 2025 · 被引用 1 次
- UniVS: Unified and Universal Video Segmentation with Prompts as QueriesMinghan Li, Shuai Li, Xindong Zhang, Lei ZhangCVPR 2024
- Explicit Visual Prompting for Low-Level Structure SegmentationsWeihuang Liu, Xi Shen, Chi-Man Pun, Xiaodong CunCVPR 2023
