Lune

CVPR2024Top-tier venue

V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Penghao Wu, Saining Xie

2024Year
32Citations
154Top-tier citations

Abstract

What color is the liquid in the glass?

"Sorry, the current visual information is not enough. I need to find the glass first."

"The glass is most likely to appear on the dining table."

"The liquid in the glass is green."

GPT4-V answer: "The liquid in the glass appears to be a shade of pink."

Figure 1. The visual search mechanism enables humans to identify a target within a multitude of stimuli, streamlining the organization of information critical for problem-solving and reasoning. In this work, we explore this core mechanism in the context of MLLMs, addressing its absence, which currently impedes precise visual grounding, especially for high-resolution images. In this example, the VQA LLM could not immediately answer the question, thus activating V * , an LLM-guided visual search process that uses common sense and contextual cues to search for the required details. Throughout this informed search, it builds a visual working memory (VWM), tokenizing the overall context and the areas of interest related to the targets, which are then re-fed to the VQA LLM, enabling it to accurately answer the question.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6fa5fe12-70b3-483e-8f36-a06b1aec6782

Cited by top-tier papers154

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines