Lune

CVPR2024顶会

V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Penghao Wu, Saining Xie

2024年份
32被引次数
154顶会引用

摘要

What color is the liquid in the glass?

"Sorry, the current visual information is not enough. I need to find the glass first."

"The glass is most likely to appear on the dining table."

"The liquid in the glass is green."

GPT4-V answer: "The liquid in the glass appears to be a shade of pink."

Figure 1. The visual search mechanism enables humans to identify a target within a multitude of stimuli, streamlining the organization of information critical for problem-solving and reasoning. In this work, we explore this core mechanism in the context of MLLMs, addressing its absence, which currently impedes precise visual grounding, especially for high-resolution images. In this example, the VQA LLM could not immediately answer the question, thus activating V * , an LLM-guided visual search process that uses common sense and contextual cues to search for the required details. Throughout this informed search, it builds a visual working memory (VWM), tokenizing the overall context and the areas of interest related to the targets, which are then re-fed to the VQA LLM, enabling it to accurately answer the question.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper154

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖