Dual-Process Image Generation
Grace Luo, Jonathan Granskog, Aleksander Holynski, Trevor Darrell
Abstract
Prior methods for controlling image generation are limited in their ability to be taught new tasks. In contrast, vision-language models, or VLMs, can learn tasks in-context and produce the correct outputs for a given input. We propose a dual-process distillation scheme that allows feed-forward image generators to learn new tasks from deliberative VLMs. Our scheme uses a VLM to rate the generated images and backpropagates this gradient to update the weights of the image generator. Our general framework enables a wide variety of new control tasks through the same text-and-image based interface. We showcase a handful of applications of this technique for different types of control signals, such as commonsense inferences and visual prompts. With our method, users can implement multimodal controls for properties such as color palette, line weight, horizon position, and relative depth within a matter of minutes. Project page: https://dual-process.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b41a516-165c-427f-b77b-0ffb86da4e25Cited by top-tier papers5
- Learning an Image Editing Model without Image Editing PairsNupur Kumari, Sheng-Yu Wang, Nanxuan Zhao, Yotam Nitzan et al.ICLR 2026 · 14 citations
- Product of Experts for Visual GenerationYunzhi Zhang, Carson Murtuza-Lanier, Zizhang Li, Yilun Du et al.ICLR 2026 · 7 citations
- PhyCo: Learning Controllable Physical Priors for Generative MotionSriram Narayanan, Ziyu Jiang, Srinivasa G. Narasimhan, Manmohan ChandrakerCVPR 2026 · 6 citations
- ReasonX: MLLM-Guided Intrinsic Image DecompositionAlara Dirik, Tuanfeng Yang Wang, Duygu Ceylan, Stefanos Zafeiriou et al.CVPR 2026 · 5 citations
- Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction ScenesWenxuan Peng, Bharath Hariharan, Hadar Averbuch-ElorSIGGRAPH 2026
Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen et al.CVPR 2026 · 124 citations
- The Narrow Gate: Localized Image-Text Communication in Native Multimodal ModelsAlessandro Serra, Francesco Ortu, Emanuele Panizon, Lucrezia Valeriani et al.NeurIPS 2025 · 4 citations
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus et al.ICML 2023 · 36 citations
- Seeing Through Words: Controlling Visual Retrieval Quality with Language ModelsJianglin Lu, Simon Jenni, Kushal Kafle, Jing Shi et al.ICLR 2026 · 3 citations
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 1 citation
