Self-Reference Image Super-Resolution via Pre-trained Diffusion Large Model and Window Adjustable Transformer
Guangyuan Li, Wei Xing, Lei Zhao, Zehua Lan, Jiakai Sun, Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, Zhijie Lin
Abstract
Currently, reference-based super-resolution (RefSR) techniques leverage high-resolution (HR) reference images to provide useful content and texture information for low-resolution (LR) images during the super-resolution (SR) process. Nevertheless, it is time-consuming, laborious, and even impossible in some cases to find high-quality reference images. To tackle this problem, we propose a brand-new self-reference image super-resolution approach using a pre-trained diffusion large model and a window adjustable transformer, termed DWTrans. Our proposed method does not require explicitly inputting manually acquired reference images during training and inference. Specifically, we feed the degraded LR images into a pre-trained stable diffusion large model to automatically generate corresponding high-quality self-reference (SRef) images that provide valuable high-frequency details for the LR images in the process of SR. To extract valuable high-frequency information in SRef images, we design a window adjustable transformer with both non-adjustable window layer (NWL) and adjustable window layer (AWL). The NWL learns local features from LR images using a dense window, while the AWL acquires global features from the SRef images using a random sparse window. Furthermore, to fully utilize the high-frequency features in the SRef image, we introduce the adaptive deformable fusion module to adaptively fuse the features of the LR and SRef images. Experimental results validate that our proposed DWTrans outperforms state-of-the-art methods on various benchmark datasets both quantitatively and visually.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get efdd4e11-82f0-4845-8f7c-b55db53a61aaCited by top-tier papers2
- CoSeR: Bridging Image and Language for Cognitive Super-ResolutionHaoze Sun, Wenbo Li, Jianzhuang Liu, Haoyu Chen et al.CVPR 2024 · 43 citations
- Cascaded Diffusion Models for Virtual Try-On: Improving Control and ResolutionGuangyuan Li, Yongkang Wang, Junsheng Luan, Lei Zhao et al.AAAI 2025 · 4 citations
Related papers
- CDFormer: When Degradation Prediction Embraces Diffusion Model for Blind Image Super-ResolutionQingguo Liu, Chenyi Zhuang, Pan Gao, Jie QinCVPR 2024 · 19 citations
- Learning Texture Transformer Network for Image Super-ResolutionFuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu et al.CVPR 2020
- Effective Diffusion Transformer Architecture for Image Super-ResolutionKun Cheng, Lei Yu, Zhijun Tu, Xiao He et al.AAAI 2025 · 26 citations
- Diffusion Transformer Meets Multi-Level Wavelet Spectrum for Single Image Super-ResolutionPeng Du, Hui Li, Han Xu, Paul Barom Jeon et al.ICCV 2025 · 3 citations
- Rethinking Diffusion Model for Multi-Contrast MRI Super-ResolutionGuangyuan Li, Chen Rao, Juncheng Mo, Zhanjie Zhang et al.CVPR 2024
