Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu, Rongyuan Wu, Lingchen Sun, Yuhui Wu, Lei Zhang
摘要
Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggressive resolution reduction of the (e.g., downsampling) in the SD model. One solution is to employ a VAE with a lower downsampling rate for diffusion; however, adapting its latent features with the pretrained UNet while mitigating the increased computational cost poses new challenges. To address these issues, we propose a Transfer VAE Training (TVT) strategy to transfer the downsampled VAE into a one while adapting to the pre-trained UNet. Specifically, we first train a decoder based on the output features of the original VAE encoder, then train a encoder while keeping the newly trained decoder fixed. Such a TVT strategy aligns the new encoder-decoder pair with the original VAE latent space while enhancing image fine details. Additionally, we introduce a compact VAE and compute-efficient UNet by optimizing their network architectures, reducing the computational cost while capturing high-resolution fine-scale features. Experimental results demonstrate that our TVT method significantly improves fine-structure preservation, which is often compromised by other SD-based methods, while requiring fewer FLOPs than state-of-the-art onestep diffusion models. The official code can be found at https://github.com/Joyies/TVT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- One-Step Diffusion Transformer for Controllable Real-World Image Super-ResolutionYushun Fang, Yuxiang Chen, Shibo Yin, Qiang Hu 等CVPR 2026 · 被引用 9 次
- GenDR: Lighten Generative Detail RestorationYan Wang, Shijie Zhao, Kexin Zhang, Junlin Li 等ICLR 2026 · 被引用 5 次
- GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-ResolutionQiaosi Yi, Shuai Li, Rongyuan Wu, Lingchen Sun 等CVPR 2026 · 被引用 4 次
- Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-ResolutionHao Chen, Junyang Chen, Jinshan Pan, Jiangxin DongCVPR 2026 · 被引用 4 次
- InstructRestore: Region-Customized Image Restoration with Human InstructionsShuaizheng Liu, Jianqi Ma, Lingchen Sun, Xiangtao Kong 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper45
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- DA-VAE: Plug-in Latent Compression for Diffusion via Detail AlignmentXin Cai, Zhiyuan You, Zhoutong Zhang, Tianfan XueCVPR 2026 · 被引用 3 次
- Time-Aware One Step Diffusion Network for Real-World Image Super-ResolutionTianyi Zhang, Zheng-Peng Duan, Chun-Le Guo, Peng-Tao Jiang 等CVPR 2026 · 被引用 6 次
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 被引用 319 次
- WEVSR: Video Diffusion Generators for Real-World Video Super‑Resolution with Wavelet-Enhanced VAE EncoderYuying Chen, Liu, Linyan Jiang, Qifan Gao 等ICML 2026
- TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-ResolutionLinwei Dong, Qingnan Fan, Yihong Guo, Zhonghao Wang 等CVPR 2025
