Lune

NeurIPS2025Top-tier venue

Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation

Bailey Trang Nguyen, Parham Saremi, Alan Q. Wang, Fangrui Huang, Zahra Tehraninasab, Amar Kumar, Tal Arbel, Fei-Fei Li, Ehsan Adeli

2025Year
2Citations

Abstract

Capturing diversity is crucial in conditional and prompt-based image generation, particularly when conditions contain uncertainty that can lead to multiple plausible outputs. To generate diverse images reflecting this diversity, traditional methods often modify random seeds, making it difficult to discern meaningful differences between samples, or diversify the input prompt, which is limited in verbally interpretable diversity. We propose Rainbow, a novel conditional image generation framework, applicable to any pretrained conditional generative model, that addresses inherent condition/prompt uncertainty and generates diverse plausible images. Rainbow is based on a simple yet effective idea: decomposing the input condition into diverse latent representations, each capturing an aspect of the uncertainty and generating a distinct image. First, we integrate a latent graph, parameterized by Generative Flow Networks (GFlowNets), into the prompt representation computation. Second, leveraging GFlowNets' advanced graph sampling capabilities to capture uncertainty and output diverse trajectories over the graph, we produce multiple trajectories that collectively represent the input condition, leading to diverse condition representations and corresponding output images. Evaluations on natural image and medical image datasets demonstrate Rainbow's improvement in both diversity and fidelity across image synthesis, image generation, and counterfactual generation tasks.

  • Corresponding author. 2 We refer to "prompts" and "conditions" interchangeably.

39th Conference on Neural Information Processing Systems (NeurIPS 2025).

despite having identical conditions due to subject-specific and medical scanner-specific details. In both cases, failing to address inherent uncertainty and capturing diversity in generative models can lead to suboptimal decision-making, misinterpretations, and generation collapse, where limited and uniform outputs fail to represent the necessary variability [13,18,38,39,58,73].

Previous attempts at generating diverse images in conditional image generation models, such as using GANs [19], diffusion models [10,24], and latent diffusion models [62] (LDMs), can be categorized into two main approaches: (1) Traditional methods typically rely on randomness; for example, repeating the generation process with different random seeds or varying the random noise on the same seeds in diffusion models [25,34,49,51] to create multiple outputs. While these methods can produce non-identical images, they often fail to capture true diversity of choices and may exhibit inherent biases; (2) Another line of work involves diversifying and adding details to the input prompt verbally using a pretrained Large Language Model (LLM) such as ChatGPT [22,60,82]. Although this approach can enhance the richness of the generated content, it is confined to text-based conditions and relies on external LLM models and their own biases. Consequently, these strategies may not adequately address the inherent uncertainty of conditional image generation tasks. In addition, a more versatile approach is needed to handle multiple condition types. For instance, generating medical images conditioned on age, sex, diseases, or other medical details can enrich datasets in fields where data collection is costly and time-consuming, such as in 3D brain MRI or chest X-ray datasets.

Addressing these limitations, we introduce Rainbow, a novel conditional image generation framework designed to produce diverse and plausible images. Rainbow can be integrated into any pretrained conditional image generative model. The primary idea is to create multiple images simultaneously that capture uncertainty by collectively reflecting the input condition. To achieve this, we aim to generate diverse condition representations that encapsulate various aspects of the uncertainty inherent in the input condition within the latent space. Each representation produces a distinct output image while the pretrained generative models remain frozen or minimally modified. As a result, Rainbow delivers a range of high-quality images that comprehensively interpret the input prompt.

To achieve diverse condition latent representations that collectively reflect the input condition, we first construct a graph structure, called the latent graph, within the latent representation computation. Next, we utilize Generative Flow Networks (GFlowNets) [4,5] to sample diverse trajectories over the graph collectively representing the input condition. Specifically, GFlowNets are designed to capture uncertainty in tasks with multiple possible outputs (multiple modes) by sampling diverse high-quality intermediate representations (e.g., trajectories over a graph) that lead to varied outputs, each representing one possible optimal outcome (one mode of the solution space). GFlowNets have been applied to many contexts, including molecule generation [4], gene regulatory networks [3,48], and dropout masks [

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines