Right this way: Can VLMs Guide Us to See More to Answer Questions?
Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti, Lei Ding, Yi Zhang, Leilani Gilpin
Abstract
In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the sufficiency of the information. To investigate this gap, we identify a critical and challenging task in the Visual Question Answering (VQA) scenario: can VLMs indicate how to adjust an image when the visual information is insufficient to answer a question? This capability is especially valuable for assisting visually impaired individuals who often need guidance to capture images correctly. To evaluate this capability of current VLMs, we introduce a human-labeled dataset as a benchmark for this task. Additionally, we present an automated framework that generates synthetic training data by simulating ``where to know'' scenarios. Our empirical results show significant performance improvements in mainstream VLMs when fine-tuned with this synthetic data. This study demonstrates the potential to narrow the gap between information assessment and acquisition in VLMs, bringing their performance closer to humans.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02e3c82c-e580-4e26-9014-24b29cbdb377Cited by top-tier papers4
- Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-PracticingDiji Yang, Linda Zeng, Jinmeng Rao, Yi ZhangSIGIR 2025 · 13 citations
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
- Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind UsersFarnaz Zamiri Zeraati, Yang Trista Cao, Yuehan Qiao, Hal Daumé III et al.CHI 2026 · 1 citation
- ViBR: Automated Bug Replay from Video-Based Reports using Vision-Language ModelsSidong Feng, Dingbang Wang, Nikola Tomic, Tingting Yu et al.FSE 2026 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
Related papers
- WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelZhiqiang Yuan, Ting Zhang, Yeshuang Zhu, Jiapei Zhang et al.ICCV 2025 · 3 citations
- From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data SynthesisChuanqi Cheng, Jian Guan, Wei Wu, Rui YanEMNLP 2024 · 1 citation
- AQuA: Toward Strategic Response Generation for Ambiguous Visual QuestionsJihyoung Jang, Hyounghun KimICLR 2026 · 1 citation
- CommVQA: Situating Visual Question Answering in Communicative ContextsNandita Naik, Christopher Potts, Elisa KreissEMNLP 2024
- Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsPu Jian, Donglei Yu, Wen Yang, Shuo Ren et al.ACL 2025
