Learned Representation-Guided Diffusion Models for Large-Image Generation
Alexandros Graikos, Srikar Yellapragada, Minh-Quan Le, Saarthak Kapse, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras
摘要
To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed by domain experts and involves hundreds of millions of patches. Modern-day self-supervised learning (SSL) representations encode rich semantic and visual information. In this paper, we posit that such representations are expressive enough to act as proxies to finegrained human labels. We introduce a novel approach that trains diffusion models conditioned on embeddings from SSL. Our diffusion models successfully project these features back to high-quality histopathology and remote sensing images. In addition, we construct larger images by assembling spatially consistent patches inferred from SSL embeddings, preserving long-range dependencies. Augmenting real data by generating variations of real images improves downstream classifier accuracy for patch-level and larger, image-scale classification tasks. Our models are effective even on datasets not encountered during training, demonstrating their robustness and generalizability. Generating images from learned embeddings is agnostic to the source of the embeddings. The SSL embeddings used to generate a large image can either be extracted from a reference image, or sampled from an auxiliary model conditioned on any related modality (e.g. class labels, text, genomic data). As proof of concept, we introduce the text-to-large image synthesis paradigm where we successfully synthesize large pathology and satellite images out of text descriptions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingZhihao Huang, Xi Qiu, Yukuo Ma, Yifu Zhou 等NeurIPS 2025 · 被引用 20 次
- PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask ConditionsMahesh Bhosale, Abdul Wasi, Yuanhao Zhai, Yunjie Tian 等ICCV 2025 · 被引用 8 次
- Importance-Based Token Merging for Efficient Image and Video GenerationHaoyu Wu, Jingyi Xu, Hieu Le, Dimitris SamarasICCV 2025 · 被引用 3 次
- G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction Via Evolutionary DiffusionMengdi Liu, Zhangyang Gao, Hong Chang, Stan Z. Li 等ICCV 2025 · 被引用 2 次
- D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide ImagesShurong Yang, Dong Wei, Yihuang Hu, Qiong Peng 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- ZoomLDM: Latent Diffusion Model for Multi-scale Image GenerationSrikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Prateek Prasanna 等CVPR 2025
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion ModelsSamuel Lavoie, Michael Noukhovitch, Aaron C. CourvilleNeurIPS 2025 · 被引用 3 次
- Dynamic Entity-Masked Graph Diffusion Model for Histopathology Image Representation LearningZhenfeng Zhuang, Min Cen, Yanfeng Li, Fangyu Zhou 等AAAI 2025
- Can Generative Models Improve Self-Supervised Representation Learning?Sana Ayromlou, Vahid Reza Khazaie, Fereshteh Forghani, Arash AfkanpourAAAI 2025 · 被引用 5 次
- Semantic and Visual Crop-Guided Diffusion Models for Heterogeneous Tissue Synthesis in HistopathologySaghir Alfasly, Wataru Uegami, Md. Enamul Hoq, Ghazal Alabtah 等NeurIPS 2025 · 被引用 3 次
