Lune

NeurIPS2025顶会

RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation

Boyuan Cao, Jiaxin Ye, Yujie Wei, Hongming Shan

2025年份
10被引次数
6顶会引用

摘要

attention to improve the structural consistency of the latent representation towards high-quality images at the training resolution. (iii) We propose progressively upsampling the resolution of latent representation in the pixel space, which can alleviate the artifacts caused by the latent space upsampling. (iv) Extensive experimental results demonstrate that the proposed RepLDM significantly outperforms the SOTA models in terms of image quality and inference time, emphasizing its great potential for real-world applications.

2 Related Work HR image generation with super-resolution. An intuitive approach to generating HR images is to first use a pre-trained LDM to generate training-resolution 3 (TR) images and then apply a superresolution model to perform upsampling [26,31,47,48,54]. Although one can obtain structurally consistent HR images in this way, super-resolution models are primarily focused on enlarging the image, and shown to be unable to produce the details that users expect in HR images [6,27,28].

Existing additional training methods either fine-tune existing LDMs with HR images [10,19,57] or train cascaded diffusion models to gradually synthesize higher-resolution images [17,44]. Though effective, these methods require expensive training resources that are unaffordable for regular users.

HR image generation in training-free manner. Current training-free methods can be roughly classified into three categories: sliding window-based, parameter rectification-based, and progressive upsampling-based methods. Sliding window-based methods consider spatially splitting HR image generation [1,12,25]. Specifically, they partition an HR image into several patches with overlap, and then denoise each patch. However, due to the lack of communication between windows, these methods result in structural disarray and content duplication. While enlarging the overlaps of the windows mitigates this issue, it can result in unbearable computational costs. For the parameter rectification-based methods, some researchers discovered that the collapse of HR image generation is due to the mismatches between higher resolutions and the model's parameters [14,[20][21][22]56]. These methods attempt to eliminate the mismatches by rectifying the parameters such as the dilation rates of some convolutional layers. While mitigating the structural inconsistency, they often lead to the degradation of image details. Different from the aforementioned two types, the progressive upsampling-based methods show SOTA performance in some recent studies [6,27,28,37]. Though promising, they require fully repeating the denoising process multiple times, which incurs unbearable computational overhead. Additionally, these methods perform upsampling in the latent space, which may introduce artifacts.

Although their remarkable results, these methods fail to improve the quality of HR images and computational efficiency at the same time. In contrast, RepLDM aims to generate HR images with high quality and high efficiency, towards practical applications.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper36

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖