Generative Visual Prompt: Unifying Distributional Control of Pre-Trained Generative Models
Chen Henry Wu, Saman Motamed, Shaunak Srivastava, Fernando De la Torre
Abstract
Generative models (e.g., GANs, diffusion models) learn the underlying data distribution in an unsupervised manner. However, many applications of interest require sampling from a particular region of the output space or sampling evenly over a range of characteristics. For efficient sampling in these scenarios, we propose Generative Visual Prompt (PromptGen), a framework for distributional control over pre-trained generative models by incorporating knowledge of other off-the-shelf models. PromptGen defines control as energy-based models (EBMs) and samples images in a feed-forward manner by approximating the EBM with invertible neural networks, avoiding optimization at inference. Our experiments demonstrate how PromptGen can efficiently sample from several unconditional generative models (e.g., StyleGAN2, StyleNeRF, diffusion autoencoder, NVAE) in a controlled or/and de-biased manner using various off-the-shelf models: (1) with the CLIP model as control, PromptGen can sample images guided by text, (2) with image classifiers as control, PromptGen can de-bias generative models across a set of attributes or attribute combinations, and (3) with inverse graphics models as control, PromptGen can sample images of the same identity in different poses. (4) Finally, PromptGen reveals that the CLIP model shows a"reporting bias"when used as control, and PromptGen can further de-bias this controlled distribution in an iterative manner. The code is available at https://github.com/ChenWu98/Generative-Visual-Prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0eea9704-e6a9-4518-bc19-1cd7659eff4fCited by top-tier papers10
- A Latent Space of Stochastic Diffusion Models for Zero-Shot Image Editing and GuidanceChen Henry Wu, Fernando De la TorreICCV 2023 · 141 citations
- Finetuning Text-to-Image Diffusion Models for FairnessXudong Shen, Chao Du, Tianyu Pang, Min Lin et al.ICLR 2024 · 97 citations
- ITI-Gen: Inclusive Text-to-Image GenerationCheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu et al.ICCV 2023 · 89 citations
- A Distributional Lens for Multi-Aspect Controllable Text GenerationYuxuan Gu, Xiaocheng Feng, Sicheng Ma, Lingyuan Zhang et al.EMNLP 2022 · 16 citations
- Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image DomainHanyue Lou, Jinxiu (Sherry) Liang, Minggui Teng, Bin Fan et al.NeurIPS 2024 · 13 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- InvDiff: Invariant Guidance for Bias Mitigation in Diffusion ModelsMin Hou, Yueying Wu, Chang Xu, Yu-Hao Huang et al.KDD 2025 · 2 citations
- Any-Shift Prompting for Generalization Over DistributionsZehao Xiao, Jiayi Shen, Mohammad Mahdi Derakhshani, Shengcai Liao et al.CVPR 2024 · 3 citations
- Anti-Exposure Bias in Diffusion ModelsJunyu Zhang, Daochang Liu, Eunbyung Park, Shichao Zhang et al.ICLR 2025
- Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language ModelsDonghoon Kim, Minji Bae, Kyuhong Shim, Byonghyo ShimICLR 2025
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 458 citations
