RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation
Boyuan Cao, Jiaxin Ye, Yujie Wei, Hongming Shan
Abstract
attention to improve the structural consistency of the latent representation towards high-quality images at the training resolution. (iii) We propose progressively upsampling the resolution of latent representation in the pixel space, which can alleviate the artifacts caused by the latent space upsampling. (iv) Extensive experimental results demonstrate that the proposed RepLDM significantly outperforms the SOTA models in terms of image quality and inference time, emphasizing its great potential for real-world applications.
2 Related Work HR image generation with super-resolution. An intuitive approach to generating HR images is to first use a pre-trained LDM to generate training-resolution 3 (TR) images and then apply a superresolution model to perform upsampling [26,31,47,48,54]. Although one can obtain structurally consistent HR images in this way, super-resolution models are primarily focused on enlarging the image, and shown to be unable to produce the details that users expect in HR images [6,27,28].
Existing additional training methods either fine-tune existing LDMs with HR images [10,19,57] or train cascaded diffusion models to gradually synthesize higher-resolution images [17,44]. Though effective, these methods require expensive training resources that are unaffordable for regular users.
HR image generation in training-free manner. Current training-free methods can be roughly classified into three categories: sliding window-based, parameter rectification-based, and progressive upsampling-based methods. Sliding window-based methods consider spatially splitting HR image generation [1,12,25]. Specifically, they partition an HR image into several patches with overlap, and then denoise each patch. However, due to the lack of communication between windows, these methods result in structural disarray and content duplication. While enlarging the overlaps of the windows mitigates this issue, it can result in unbearable computational costs. For the parameter rectification-based methods, some researchers discovered that the collapse of HR image generation is due to the mismatches between higher resolutions and the model's parameters [14,[20][21][22]56]. These methods attempt to eliminate the mismatches by rectifying the parameters such as the dilation rates of some convolutional layers. While mitigating the structural inconsistency, they often lead to the degradation of image details. Different from the aforementioned two types, the progressive upsampling-based methods show SOTA performance in some recent studies [6,27,28,37]. Though promising, they require fully repeating the denoising process multiple times, which incurs unbearable computational overhead. Additionally, these methods perform upsampling in the latent space, which may introduce artifacts.
Although their remarkable results, these methods fail to improve the quality of HR images and computational efficiency at the same time. In contrast, RepLDM aims to generate HR images with high quality and high efficiency, towards practical applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0dde1d0-6960-4ecd-add4-a6a6f74238d7Cited by top-tier papers6
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image GenerationSashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao et al.CVPR 2026 · 7 citations
- FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale FusionHaonan Qiu, Shiwei Zhang, Yujie Wei, Ruihang Chu et al.ICCV 2025 · 5 citations
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen et al.CVPR 2026 · 3 citations
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao et al.CVPR 2026 · 1 citation
- S²Flow: Towards Fast and Authentic Training-Free High-Resolution Video GenerationChaoqun Wang, Shaobo Min, Xu YangAAAI 2026
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolutionJunyang Chen, Jinshan Pan, Jiangxin DongCVPR 2025
- PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and GenerationLiyao Jiang, Negar Hassanpour, Mohammad Salameh, Mohammadreza Samadi et al.AAAI 2025 · 8 citations
- DiffuseHigh: Training-Free Progressive High-Resolution Image Synthesis Through Structure GuidanceYounghyun Kim, Geunmin Hwang, Junyu Zhang, Eunbyung ParkAAAI 2025 · 30 citations
- ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained GuidanceShuwei Shi, Wenbo Li, Yuechen Zhang, Jingwen He et al.AAAI 2025 · 23 citations
