Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang
Abstract
User prompts for generative AI models are often underspecified, leading to a misalignment between the user intent and models' understanding. As a result, users commonly have to painstakingly refine their prompts. We study this alignment problem in text-to-image (T2I) generation and propose a prototype for proactive T2I agents equipped with an interface to (1) actively ask clarification questions when uncertain, and (2) present their uncertainty about user intent as an understandable and editable belief graph. We build simple prototypes for such agents and propose a new scalable and automated evaluation approach using two agents, one with a ground truth intent (an image) while the other tries to ask as few questions as possible to align with the ground truth. We experiment over three image-text datasets: ImageInWords (Garg et al., 2024), COCO (Lin et al., 2014) and DesignBench, a benchmark we curated with strong artistic and design elements. Experiments over the three datasets demonstrate the proposed T2I agents' ability to ask informative questions and elicit crucial information to achieve successful alignment with at least 2 times higher VQAScore (Lin et al., 2024) than the standard T2I generation. Moreover, we conducted human studies and observed that at least 90% of human subjects found these agents and their belief graphs helpful for their T2I workflow, highlighting the effectiveness of our approach. Code and DesignBench can be found at https: //github.com/google-deepmind/ proactive_t2i_agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi et al.NeurIPS 2025 · 12 citations
- Charts Are Not Images: On the Challenges of Scientific Chart EditingShawn Li, Ryan Rossi, Sungchul Kim, Sunav Choudhary et al.ICLR 2026 · 11 citations
- Twin Co-Adaptive Dialogue for Progressive Image GenerationJianhui Wang, Yangfan He, Yan Zhong, Xinyuan Song et al.ACM MM 2025 · 4 citations
- Morae: Proactively Pausing UI Agents for User ChoicesYi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy PavelUIST 2025 · 4 citations
- Safe Autoregressive Image Generation with Iterative Self-Improving CodebooksYunqi Xue, Zhijiang Li, Phil Torr, Jindong GuICML 2026 · 1 citation
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
- Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to DesignQian Yang, Aaron Steinfeld, Carolyn P. Rosé, John ZimmermanCHI 2020 · 604 citations
Related papers
- Resolving Ambiguities in Text-to-Image Generative ModelsNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala et al.ACL 2023 · 8 citations
- LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured AnnotationsZhichao Yang, Tianjiao Gu, Jianjie Wang, Feiyu Lin et al.AAAI 2026 · 1 citation
- Is It AI or Is It Me? Understanding Users' Prompt Journey with Text-to-Image Generative AI ToolsAtefeh Mahdavi Goloujeh, Anne Sullivan, Brian MagerkoCHI 2024 · 88 citations
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationChieh-Yun Chen, Min Shi, Gong Zhang, Humphrey ShiICCV 2025 · 1 citation
- Bridging Gulfs in UI Generation through Semantic GuidanceSeokhyeon Park, Soohyun Lee, Eugene Choi, Hyunwoo Kim et al.CHI 2026 · 3 citations
