RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation
Boyuan Cao, Jiaxin Ye, Yujie Wei, Hongming Shan
摘要
attention to improve the structural consistency of the latent representation towards high-quality images at the training resolution. (iii) We propose progressively upsampling the resolution of latent representation in the pixel space, which can alleviate the artifacts caused by the latent space upsampling. (iv) Extensive experimental results demonstrate that the proposed RepLDM significantly outperforms the SOTA models in terms of image quality and inference time, emphasizing its great potential for real-world applications.
2 Related Work HR image generation with super-resolution. An intuitive approach to generating HR images is to first use a pre-trained LDM to generate training-resolution 3 (TR) images and then apply a superresolution model to perform upsampling [26,31,47,48,54]. Although one can obtain structurally consistent HR images in this way, super-resolution models are primarily focused on enlarging the image, and shown to be unable to produce the details that users expect in HR images [6,27,28].
Existing additional training methods either fine-tune existing LDMs with HR images [10,19,57] or train cascaded diffusion models to gradually synthesize higher-resolution images [17,44]. Though effective, these methods require expensive training resources that are unaffordable for regular users.
HR image generation in training-free manner. Current training-free methods can be roughly classified into three categories: sliding window-based, parameter rectification-based, and progressive upsampling-based methods. Sliding window-based methods consider spatially splitting HR image generation [1,12,25]. Specifically, they partition an HR image into several patches with overlap, and then denoise each patch. However, due to the lack of communication between windows, these methods result in structural disarray and content duplication. While enlarging the overlaps of the windows mitigates this issue, it can result in unbearable computational costs. For the parameter rectification-based methods, some researchers discovered that the collapse of HR image generation is due to the mismatches between higher resolutions and the model's parameters [14,[20][21][22]56]. These methods attempt to eliminate the mismatches by rectifying the parameters such as the dilation rates of some convolutional layers. While mitigating the structural inconsistency, they often lead to the degradation of image details. Different from the aforementioned two types, the progressive upsampling-based methods show SOTA performance in some recent studies [6,27,28,37]. Though promising, they require fully repeating the denoising process multiple times, which incurs unbearable computational overhead. Additionally, these methods perform upsampling in the latent space, which may introduce artifacts.
Although their remarkable results, these methods fail to improve the quality of HR images and computational efficiency at the same time. In contrast, RepLDM aims to generate HR images with high quality and high efficiency, towards practical applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image GenerationSashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao 等CVPR 2026 · 被引用 7 次
- FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale FusionHaonan Qiu, Shiwei Zhang, Yujie Wei, Ruihang Chu 等ICCV 2025 · 被引用 5 次
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen 等CVPR 2026 · 被引用 3 次
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao 等CVPR 2026 · 被引用 1 次
- S²Flow: Towards Fast and Authentic Training-Free High-Resolution Video GenerationChaoqun Wang, Shaobo Min, Xu YangAAAI 2026
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
相关 Paper
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolutionJunyang Chen, Jinshan Pan, Jiangxin DongCVPR 2025
- PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and GenerationLiyao Jiang, Negar Hassanpour, Mohammad Salameh, Mohammadreza Samadi 等AAAI 2025 · 被引用 8 次
- DiffuseHigh: Training-Free Progressive High-Resolution Image Synthesis Through Structure GuidanceYounghyun Kim, Geunmin Hwang, Junyu Zhang, Eunbyung ParkAAAI 2025 · 被引用 30 次
- ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained GuidanceShuwei Shi, Wenbo Li, Yuechen Zhang, Jingwen He 等AAAI 2025 · 被引用 23 次
