ZoomLDM: Latent Diffusion Model for Multi-scale Image Generation
Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Prateek Prasanna, Rajarsi Gupta, Joel H. Saltz, Dimitris Samaras
Abstract
Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Given that it is infeasible to directly train a model on 'whole' images from domains with potential gigapixel sizes, diffusion-based generative methods have focused on synthesizing small, fixed-size patches extracted from these images. However, generating small patches has limited applicability since patch-based models fail to capture the global structures and wider context of large images, which can be crucial for synthesizing (semantically) accurate samples. To overcome this limitation, we present ZoomLDM, a diffusion model tailored for generating images across multiple scales. Central to our approach is a novel magnificationaware conditioning mechanism that utilizes self-supervised learning (SSL) embeddings and allows the diffusion model to synthesize images at different 'zoom' levels, i.e., fixedsize patches extracted from large images at varying scales. ZoomLDM synthesizes coherent histopathology images that remain contextually accurate and detailed at different zoom levels, achieving state-of-the-art image generation quality across all scales and excelling in the data-scarce setting of generating thumbnails of entire large images. The multi-scale nature of ZoomLDM unlocks additional capabilities in large image generation, enabling computationally tractable and globally coherent image synthesis up to 4096 × 4096 pixels and 4× super-resolution. Additionally, multi-scale features extracted from ZoomLDM are highly effective in multiple instance learning experiments. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c35f04a8-1766-4b41-9a76-cebad82b4dd8Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Learned Representation-Guided Diffusion Models for Large-Image GenerationAlexandros Graikos, Srikar Yellapragada, Minh-Quan Le, Saarthak Kapse et al.CVPR 2024
- Generative Powers of TenXiaojuan Wang, Janne Kontkanen, Brian Curless, Steven M. Seitz et al.CVPR 2024 · 3 citations
- Semantic and Visual Crop-Guided Diffusion Models for Heterogeneous Tissue Synthesis in HistopathologySaghir Alfasly, Wataru Uegami, Md. Enamul Hoq, Ghazal Alabtah et al.NeurIPS 2025 · 3 citations
- DiffuseHigh: Training-Free Progressive High-Resolution Image Synthesis Through Structure GuidanceYounghyun Kim, Geunmin Hwang, Junyu Zhang, Eunbyung ParkAAAI 2025 · 30 citations
- DiffusionSat: A Generative Foundation Model for Satellite ImagerySamar Khanna, Patrick Liu, Linqi Zhou, Chenlin Meng et al.ICLR 2024 · 173 citations
