Learned Representation-Guided Diffusion Models for Large-Image Generation
Alexandros Graikos, Srikar Yellapragada, Minh-Quan Le, Saarthak Kapse, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras
Abstract
To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed by domain experts and involves hundreds of millions of patches. Modern-day self-supervised learning (SSL) representations encode rich semantic and visual information. In this paper, we posit that such representations are expressive enough to act as proxies to finegrained human labels. We introduce a novel approach that trains diffusion models conditioned on embeddings from SSL. Our diffusion models successfully project these features back to high-quality histopathology and remote sensing images. In addition, we construct larger images by assembling spatially consistent patches inferred from SSL embeddings, preserving long-range dependencies. Augmenting real data by generating variations of real images improves downstream classifier accuracy for patch-level and larger, image-scale classification tasks. Our models are effective even on datasets not encountered during training, demonstrating their robustness and generalizability. Generating images from learned embeddings is agnostic to the source of the embeddings. The SSL embeddings used to generate a large image can either be extracted from a reference image, or sampled from an auxiliary model conditioned on any related modality (e.g. class labels, text, genomic data). As proof of concept, we introduce the text-to-large image synthesis paradigm where we successfully synthesize large pathology and satellite images out of text descriptions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a796bd2-9c45-4a59-9636-571630874582Cited by top-tier papers11
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingZhihao Huang, Xi Qiu, Yukuo Ma, Yifu Zhou et al.NeurIPS 2025 · 20 citations
- PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask ConditionsMahesh Bhosale, Abdul Wasi, Yuanhao Zhai, Yunjie Tian et al.ICCV 2025 · 8 citations
- Importance-Based Token Merging for Efficient Image and Video GenerationHaoyu Wu, Jingyi Xu, Hieu Le, Dimitris SamarasICCV 2025 · 3 citations
- G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction Via Evolutionary DiffusionMengdi Liu, Zhangyang Gao, Hong Chang, Stan Z. Li et al.ICCV 2025 · 2 citations
- D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide ImagesShurong Yang, Dong Wei, Yihuang Hu, Qiong Peng et al.NeurIPS 2025 · 2 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- ZoomLDM: Latent Diffusion Model for Multi-scale Image GenerationSrikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Prateek Prasanna et al.CVPR 2025
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion ModelsSamuel Lavoie, Michael Noukhovitch, Aaron C. CourvilleNeurIPS 2025 · 3 citations
- Dynamic Entity-Masked Graph Diffusion Model for Histopathology Image Representation LearningZhenfeng Zhuang, Min Cen, Yanfeng Li, Fangyu Zhou et al.AAAI 2025
- Can Generative Models Improve Self-Supervised Representation Learning?Sana Ayromlou, Vahid Reza Khazaie, Fereshteh Forghani, Arash AfkanpourAAAI 2025 · 5 citations
- Semantic and Visual Crop-Guided Diffusion Models for Heterogeneous Tissue Synthesis in HistopathologySaghir Alfasly, Wataru Uegami, Md. Enamul Hoq, Ghazal Alabtah et al.NeurIPS 2025 · 3 citations
