Lune

NeurIPS2025顶会

Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation

Bailey Trang Nguyen, Parham Saremi, Alan Q. Wang, Fangrui Huang, Zahra Tehraninasab, Amar Kumar, Tal Arbel, Fei-Fei Li, Ehsan Adeli

2025年份
2被引次数

摘要

Capturing diversity is crucial in conditional and prompt-based image generation, particularly when conditions contain uncertainty that can lead to multiple plausible outputs. To generate diverse images reflecting this diversity, traditional methods often modify random seeds, making it difficult to discern meaningful differences between samples, or diversify the input prompt, which is limited in verbally interpretable diversity. We propose Rainbow, a novel conditional image generation framework, applicable to any pretrained conditional generative model, that addresses inherent condition/prompt uncertainty and generates diverse plausible images. Rainbow is based on a simple yet effective idea: decomposing the input condition into diverse latent representations, each capturing an aspect of the uncertainty and generating a distinct image. First, we integrate a latent graph, parameterized by Generative Flow Networks (GFlowNets), into the prompt representation computation. Second, leveraging GFlowNets' advanced graph sampling capabilities to capture uncertainty and output diverse trajectories over the graph, we produce multiple trajectories that collectively represent the input condition, leading to diverse condition representations and corresponding output images. Evaluations on natural image and medical image datasets demonstrate Rainbow's improvement in both diversity and fidelity across image synthesis, image generation, and counterfactual generation tasks.

  • Corresponding author. 2 We refer to "prompts" and "conditions" interchangeably.

39th Conference on Neural Information Processing Systems (NeurIPS 2025).

despite having identical conditions due to subject-specific and medical scanner-specific details. In both cases, failing to address inherent uncertainty and capturing diversity in generative models can lead to suboptimal decision-making, misinterpretations, and generation collapse, where limited and uniform outputs fail to represent the necessary variability [13,18,38,39,58,73].

Previous attempts at generating diverse images in conditional image generation models, such as using GANs [19], diffusion models [10,24], and latent diffusion models [62] (LDMs), can be categorized into two main approaches: (1) Traditional methods typically rely on randomness; for example, repeating the generation process with different random seeds or varying the random noise on the same seeds in diffusion models [25,34,49,51] to create multiple outputs. While these methods can produce non-identical images, they often fail to capture true diversity of choices and may exhibit inherent biases; (2) Another line of work involves diversifying and adding details to the input prompt verbally using a pretrained Large Language Model (LLM) such as ChatGPT [22,60,82]. Although this approach can enhance the richness of the generated content, it is confined to text-based conditions and relies on external LLM models and their own biases. Consequently, these strategies may not adequately address the inherent uncertainty of conditional image generation tasks. In addition, a more versatile approach is needed to handle multiple condition types. For instance, generating medical images conditioned on age, sex, diseases, or other medical details can enrich datasets in fields where data collection is costly and time-consuming, such as in 3D brain MRI or chest X-ray datasets.

Addressing these limitations, we introduce Rainbow, a novel conditional image generation framework designed to produce diverse and plausible images. Rainbow can be integrated into any pretrained conditional image generative model. The primary idea is to create multiple images simultaneously that capture uncertainty by collectively reflecting the input condition. To achieve this, we aim to generate diverse condition representations that encapsulate various aspects of the uncertainty inherent in the input condition within the latent space. Each representation produces a distinct output image while the pretrained generative models remain frozen or minimally modified. As a result, Rainbow delivers a range of high-quality images that comprehensively interpret the input prompt.

To achieve diverse condition latent representations that collectively reflect the input condition, we first construct a graph structure, called the latent graph, within the latent representation computation. Next, we utilize Generative Flow Networks (GFlowNets) [4,5] to sample diverse trajectories over the graph collectively representing the input condition. Specifically, GFlowNets are designed to capture uncertainty in tasks with multiple possible outputs (multiple modes) by sampling diverse high-quality intermediate representations (e.g., trajectories over a graph) that lead to varied outputs, each representing one possible optimal outcome (one mode of the solution space). GFlowNets have been applied to many contexts, including molecule generation [4], gene regulatory networks [3,48], and dropout masks [

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖