Resolving Ambiguities in Text-to-Image Generative Models
Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard S. Zemel, Aram Galstyan, Rahul Gupta
Abstract
Natural language often contains ambiguities that can lead to misinterpretation and miscommunication. While humans can handle ambiguities effectively by asking clarifying questions and/or relying on contextual cues and commonsense knowledge, resolving ambiguities can be notoriously hard for machines. In this work, we study ambiguities that arise in text-to-image generative models. We curate the Text-toimage Ambiguity Benchmark (TAB) dataset to study different types of ambiguities in text-toimage generative models. 1 We then propose the Text-to-ImagE Disambiguation (TIED) framework to disambiguate the prompts given to the text-to-image generative models by soliciting clarifications from the end user. Through automatic and human evaluations, we show the effectiveness of our framework in generating more faithful images aligned with end user intention in the presence of ambiguities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c61eb5c4-6976-4d68-9644-e478f58f00dcCited by top-tier papers6
- FairQueue: Rethinking Prompt Learning for Fair Text-to-Image GenerationChristopher T. H. Teo, Milad Abdollahzadeh, Xinda Ma, Ngai-Man CheungNeurIPS 2024 · 8 citations
- Twin Co-Adaptive Dialogue for Progressive Image GenerationJianhui Wang, Yangfan He, Yan Zhong, Xinyuan Song et al.ACM MM 2025 · 4 citations
- Acknowledging Focus Ambiguity in Visual QuestionsChongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh et al.ICCV 2025 · 1 citation
- Automating UI Optimization through Multi-Agentic ReasoningZhipeng Li, Christoph Gebhardt, Yi-Chi Liao, Christian HolzCHI 2026 · 1 citation
- FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image EditingYufan Ren, Zicong Jiang, Tong Zhang, Søren Forchhammer et al.CVPR 2025
Builds on7
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
Related papers
- Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion GenerationSung-Lin Tsai, Bo-Lun Huang, Yu-Ting Shen, Cheng-Yu Yeo et al.ACM MM 2025 · 3 citations
- Proactive Agents for Multi-Turn Text-to-Image Generation Under UncertaintyMeera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt et al.ICML 2025
- Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQAElias Stengel-Eskin, Jimena Guallar-Blasco, Yi Zhou, Benjamin Van DurmeACL 2023 · 4 citations
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu et al.ICML 2026
- Image Generation from Contextually-Contradictory PromptsSaar Huberman, Or Patashnik, Omer Dahary, Ron Mokady et al.CVPR 2026 · 11 citations
