Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu, Rongyuan Wu, Lingchen Sun, Yuhui Wu, Lei Zhang
Abstract
Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggressive resolution reduction of the (e.g., downsampling) in the SD model. One solution is to employ a VAE with a lower downsampling rate for diffusion; however, adapting its latent features with the pretrained UNet while mitigating the increased computational cost poses new challenges. To address these issues, we propose a Transfer VAE Training (TVT) strategy to transfer the downsampled VAE into a one while adapting to the pre-trained UNet. Specifically, we first train a decoder based on the output features of the original VAE encoder, then train a encoder while keeping the newly trained decoder fixed. Such a TVT strategy aligns the new encoder-decoder pair with the original VAE latent space while enhancing image fine details. Additionally, we introduce a compact VAE and compute-efficient UNet by optimizing their network architectures, reducing the computational cost while capturing high-resolution fine-scale features. Experimental results demonstrate that our TVT method significantly improves fine-structure preservation, which is often compromised by other SD-based methods, while requiring fewer FLOPs than state-of-the-art onestep diffusion models. The official code can be found at https://github.com/Joyies/TVT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- One-Step Diffusion Transformer for Controllable Real-World Image Super-ResolutionYushun Fang, Yuxiang Chen, Shibo Yin, Qiang Hu et al.CVPR 2026 · 9 citations
- GenDR: Lighten Generative Detail RestorationYan Wang, Shijie Zhao, Kexin Zhang, Junlin Li et al.ICLR 2026 · 5 citations
- GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-ResolutionQiaosi Yi, Shuai Li, Rongyuan Wu, Lingchen Sun et al.CVPR 2026 · 4 citations
- Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-ResolutionHao Chen, Junyang Chen, Jinshan Pan, Jiangxin DongCVPR 2026 · 4 citations
- InstructRestore: Region-Customized Image Restoration with Human InstructionsShuaizheng Liu, Jianqi Ma, Lingchen Sun, Xiangtao Kong et al.NeurIPS 2025 · 3 citations
Builds on45
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- DA-VAE: Plug-in Latent Compression for Diffusion via Detail AlignmentXin Cai, Zhiyuan You, Zhoutong Zhang, Tianfan XueCVPR 2026 · 3 citations
- Time-Aware One Step Diffusion Network for Real-World Image Super-ResolutionTianyi Zhang, Zheng-Peng Duan, Chun-Le Guo, Peng-Tao Jiang et al.CVPR 2026 · 6 citations
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 319 citations
- WEVSR: Video Diffusion Generators for Real-World Video Super‑Resolution with Wavelet-Enhanced VAE EncoderYuying Chen, Liu, Linyan Jiang, Qifan Gao et al.ICML 2026
- TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-ResolutionLinwei Dong, Qingnan Fan, Yihong Guo, Zhonghao Wang et al.CVPR 2025
