Controlling Vision-Language Models for Multi-Task Image Restoration
Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao, Jens Sjölund, Thomas B. Schön
Abstract
Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates dramatically due to corrupted inputs. In this paper, we present a degradation-aware vision-language model (DA-CLIP) to better transfer pretrained vision-language models to low-level vision tasks as a multi-task framework for image restoration. More specifically, DA-CLIP trains an additional controller that adapts the fixed CLIP image encoder to predict high-quality feature embeddings. By integrating the embedding into an image restoration network via cross-attention, we are able to pilot the model to learn a high-fidelity image reconstruction. The controller itself will also output a degradation feature that matches the real corruptions of the input, yielding a natural classifier for different degradation types. In addition, we construct a mixed degradation dataset with synthetic captions for DA-CLIP training. Our approach advances state-of-the-art performance on both degradation-specific and unified image restoration tasks, showing a promising direction of prompting image restoration with large-scale pretrained vision-language models. Our code is available at https://github.com/Algolzw/daclip-uir.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82bbac9f-228f-49ba-b0f1-306a3920ff92Cited by top-tier papers50
- 4KAgent: Agentic Any Image to 4K Super-ResolutionYushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang et al.NeurIPS 2025 · 51 citations
- Debiased All-in-one Image Restoration with Task Uncertainty RegularizationGang Wu, Junjun Jiang, Yijun Wang, Kui Jiang et al.AAAI 2025 · 23 citations
- Hybrid Agents for Image RestorationBingchen Li, Xin Li, Yiting Lu, Zhibo ChenCVPR 2026 · 17 citations
- Residual Diffusion Bridge Model for Image RestorationHebaixu Wang, Jing Zhang, Haoyang Chen, Haonan Guo et al.CVPR 2026 · 16 citations
- Multi-axis Prompt and Multi-dimension Fusion Network for All-in-one Weather-degraded Image RestorationYuanbo Wen, Tao Gao, Jing Zhang, Ziqi Li et al.AAAI 2025 · 14 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Learning to Decompose Visual Features with Latent Textual PromptsFeng Wang, Manling Li, Xudong Lin, Hairong Lv et al.ICLR 2023 · 8 citations
- Improving Image Restoration Through Removing Degradations in Textual RepresentationsJingbo Lin, Zhilu Zhang, Yuxiang Wei, Dongwei Ren et al.CVPR 2024
- Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image RestorationYuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang et al.CVPR 2024
- Beyond Text: Frozen Large Language Models in Visual Signal ComprehensionLei Zhu, Fangyun Wei, Yanye LuCVPR 2024 · 12 citations
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and EnhancementWeiqi Li, Xuanyu Zhang, Bin Chen, Jingfen Xie et al.CVPR 2026 · 5 citations
