Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration
Junyuan Deng, Xinyi Wu, Yongxing Yang, Congchao Zhu, Song Wang, Zhenyao Wu
摘要
Recently, pre-trained text-to-image (T2I) models have been extensively adopted for real-world image restoration because of their powerful generative prior. However, controlling these large models for image restoration usually requires a large number of high-quality images and immense computational resources for training, which is costly and not privacy-friendly. In this paper, we find that the well-trained large T2I model (i.e., Flux) is able to produce a variety of high-quality images aligned with real-world distributions, offering an unlimited supply of training samples to mitigate the above issue. Specifically, we proposed a training data construction pipeline for image restoration, namely FluxGen, which includes unconditional image generation, image selection, and degraded image simulation. A novel light-weighted adapter (FluxIR) with squeeze-and-excitation layers is also carefully designed to control the large Diffusion Transformer (DiT)-based T2I model so that reasonable details can be restored. Experiments demonstrate that our proposed method enables the Flux model to adapt effectively to real-world image restoration tasks, achieving superior scores and visual quality on both synthetic and real-world degradation datasets - at only about 8.5% of the training cost compared to current approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video RestorationHaoran Bai, Xiaoxu Chen, Canqian Yang, Zongyao He 等ICLR 2026 · 被引用 10 次
- InstructRestore: Region-Customized Image Restoration with Human InstructionsShuaizheng Liu, Jianqi Ma, Lingchen Sun, Xiangtao Kong 等NeurIPS 2025 · 被引用 3 次
- DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion TransformerQingji Dong, Hang Dong, Mingqin Chen, Rui Zhang 等CVPR 2026 · 被引用 1 次
- Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image GenerationBaoteng Li, Xianghao Zang, Xinran Wang, Xiangyu Na 等CVPR 2026
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset CurationYuang Ai, Xiaoqiang Zhou, Huaibo Huang, Xiaotian Han 等NeurIPS 2024 · 被引用 81 次
- LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion TransformerSong Fei, Tian Ye, Lujia Wang, Lei ZhuICLR 2026 · 被引用 9 次
- DreamClean: Restoring Clean Image Using Deep Diffusion PriorJie Xiao, Ruili Feng, Han Zhang, Zhiheng Liu 等ICLR 2024 · 被引用 23 次
- Learning Diffusion Texture Priors for Image RestorationTian Ye, Sixiang Chen, Wenhao Chai, Zhaohu Xing 等CVPR 2024
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang 等ICLR 2026 · 被引用 34 次
