GenDR: Lighten Generative Detail Restoration
Yan Wang, Shijie Zhao, Kexin Zhang, Junlin Li, Li Zhang
摘要
Although recent research applying text-to-image (T2I) diffusion models to real-world super-resolution (SR) has achieved remarkable progress, the misalignment of their targets leads to a suboptimal trade-off between inference speed and detail fidelity. Specifically, the T2I task requires multiple inference steps to synthesize images matching to prompts and reduces the latent dimension to lower generating difficulty. Contrariwise, SR can restore high-frequency details in fewer inference steps, but it necessitates a more reliable variational auto-encoder (VAE) to preserve input information. However, most diffusion-based SRs are multistep and use 4-channel VAEs, while existing models with 16-channel VAEs are overqualified diffusion transformers, e.g., FLUX (12B). To align the target, we present a one-step diffusion model for generative detail restoration, GenDR, distilled from a tailored diffusion model with a larger latent space. In detail, we train a new SD2.1-VAE16 (0.9B) via representation alignment to expand the latent space without increasing the model size. Regarding step distillation, we propose consistent score identity distillation (CiD) that incorporates SR task-specific loss into score distillation to leverage more SR priors and align the training target. Furthermore, we extend CiD with adversarial learning and representation alignment (CiDA) to enhance perceptual quality and accelerate training. We also polish the pipeline to achieve a more efficient inference. Experimental results demonstrate that GenDR achieves state-of-the-art performance in both quantitative metrics and visual fidelity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
相关 Paper
- Eliminating VAE for Fast and High-Resolution Generative Detail RestorationYan Wang, Shijie Zhao, Junlin Li, Li zhangICLR 2026 · 被引用 1 次
- TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-ResolutionLinwei Dong, Qingnan Fan, Yihong Guo, Zhonghao Wang 等CVPR 2025
- One Diffusion Step to Real-World Super-Resolution via Flow Trajectory DistillationJianze Li, Jiezhang Cao, Yong Guo, Wenbo Li 等ICML 2025
- Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion DiscriminatorJianze Li, Jiezhang Cao, Zichen Zou, Xiongfei Su 等NeurIPS 2025 · 被引用 18 次
- One-Step Diffusion Distillation through Score Implicit MatchingWeijian Luo, Zemin Huang, Zhengyang Geng, J. Zico Kolter 等NeurIPS 2024 · 被引用 81 次
