DiffIR: Efficient Diffusion Model for Image Restoration
Bin Xia, Yulun Zhang, Shiyin Wang, Yitong Wang, Xinglong Wu, Yapeng Tian, Wenming Yang, Luc Van Gool
Abstract
Diffusion model (DM) has achieved SOTA performance by modeling the image synthesis process into a sequential application of a denoising network. However, different from image synthesis, image restoration (IR) has a strong constraint to generate results in accordance with ground-truth. Thus, for IR, traditional DMs running massive iterations on a large model to estimate whole images or feature maps is inefficient. To address this issue, we propose an efficient DM for IR (DiffIR), which consists of a compact IR prior extraction network (CPEN), dynamic IR transformer (DIRformer), and denoising network. Specifically, DiffIR has two training stages: pretraining and training DM. In pretraining, we input ground-truth images into CPEN S1 to capture a compact IR prior representation (IPR) to guide DIRformer. In the second stage, we train the DM to directly estimate the same IRP as pretrained CPEN S1 only using LQ images. We observe that since the IPR is only a compact vector, DiffIR can use fewer iterations than traditional DM to obtain accurate estimations and generate more stable and realistic results. Since the iterations are few, our Dif-fIR can adopt a joint optimization of CPEN S2 , DIRformer, and denoising network, which can further reduce the estimation error influence. We conduct extensive experiments on several IR tasks and achieve SOTA performance while consuming less computational costs. Code is available at https://github.com/Zj-BinXia/DiffIR . Dynamic Transformer Block (×𝑁!) Concat DownSample Conv 3×3 H×W×C Dynamic Transformer Block (×𝑁&) DownSample 𝐻 2 × 𝑊 2 ×2C 𝐻 4 × 𝑊 4 ×4C Dynamic Transformer Block (×𝑁+) DownSample Dynamic Transformer Block (×𝑁!) 𝐻 8 × 𝑊 8 ×8C Dynamic Transformer Block (×𝑁-) 𝐻 8 × 𝑊 8 ×8C UpSample 𝐻 4 × 𝑊 4 ×4C UpSample 𝐻 2 × 𝑊 2 ×2C UpSample Conv 1×1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd115657-b962-41cc-ba68-c9c1d922f54bCited by top-tier papers108
- Hierarchical Integration Diffusion Model for Realistic Image DeblurringZheng Chen, Yulun Zhang, Ding Liu, Bin Xia et al.NeurIPS 2023 · 179 citations
- Principled Probabilistic Imaging using Diffusion Models as Plug-and-Play PriorsZihui Wu, Yu Sun, Yifan Chen, Bingliang Zhang et al.NeurIPS 2024 · 128 citations
- 3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion PriorsXi Liu, Chaoyi Zhou, Siyu HuangNeurIPS 2024 · 127 citations
- Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion ModelHao Zhang, Lei Cao, Jiayi MaNeurIPS 2024 · 69 citations
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo et al.CVPR 2024 · 52 citations
Builds on33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
Related papers
- Rethinking Diffusion Model for Multi-Contrast MRI Super-ResolutionGuangyuan Li, Chen Rao, Juncheng Mo, Zhanjie Zhang et al.CVPR 2024
- Restoration based Generative ModelsJaemoo Choi, Yesom Park, Myungjoo KangICML 2023 · 5 citations
- Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion ModelChunming He, Chengyu Fang, Yulun Zhang, Longxiang Tang et al.ICLR 2025
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- CDFormer: When Degradation Prediction Embraces Diffusion Model for Blind Image Super-ResolutionQingguo Liu, Chenyi Zhuang, Pan Gao, Jie QinCVPR 2024 · 19 citations
