V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Penghao Wu, Saining Xie
Abstract
What color is the liquid in the glass?
"Sorry, the current visual information is not enough. I need to find the glass first."
"The glass is most likely to appear on the dining table."
"The liquid in the glass is green."
GPT4-V answer: "The liquid in the glass appears to be a shade of pink."
Figure 1. The visual search mechanism enables humans to identify a target within a multitude of stimuli, streamlining the organization of information critical for problem-solving and reasoning. In this work, we explore this core mechanism in the context of MLLMs, addressing its absence, which currently impedes precise visual grounding, especially for high-resolution images. In this example, the VQA LLM could not immediately answer the question, thus activating V * , an LLM-guided visual search process that uses common sense and contextual cues to search for the required details. Throughout this informed search, it builds a visual working memory (VWM), tokenizing the overall context and the areas of interest related to the targets, which are then re-fed to the VQA LLM, enabling it to accurately answer the question.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fa5fe12-70b3-483e-8f36-a06b1aec6782Cited by top-tier papers154
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang et al.ACL 2025 · 377 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao et al.ICLR 2026 · 321 citations
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang et al.NeurIPS 2025 · 151 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question AnsweringJunnan Dong, Qinggang Zhang, Huachi Zhou, Daochen Zha et al.ACL 2024 · 11 citations
- VGR: Visual Grounded ReasoningJiacong Wang, Zijian Kang, Haochen Wang, Xiao Liang et al.ICLR 2026 · 64 citations
- Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMsXuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen et al.CVPR 2026 · 1 citation
- DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language ModelsYangfu Li, Hongjian Zhan, Jiawei Chen, Yuning Gong et al.CVPR 2026 · 6 citations
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringWenlong Fang, Qiaofeng Wu, Jing Chen, Yun XueCVPR 2025
