Lune

ICML2026Top-tier venue

NAVIGATE: Evaluating Visual-Guided Search Decision-Making on the Open Web

YaoQi Fan, Zhe Chen, Zhu Wei, Kangxin Yin, Yangzhou Liu, Yue Cao, Zhi Zhu, Tong Lu

2026Year

Abstract

Vision-Language Models (VLMs) are increasingly deployed with web search tools, yet we still lack benchmarks that isolate a critical capability for real-world use: deciding when to search and how to steer search from ambiguous visual evidence, especially when multiple images provide overlapping or conflicting cues. We present NAVIGATE, a novel benchmark centered on images as primary evidence for open-web search planning and multi-step reasoning. It contains 500 questions across 20 domains and spans three difficulty tiers, from single-image, self-contained problems to multi-image joint search and multidomain composition. Unlike prior benchmarks that specify explicit search targets, NAVIGATE evaluates search decision-making: models must infer whether external search is necessary and iteratively refine search directions based on holistic reasoning over visual cues. Across a broad set of VLMs and search-enabled systems, performance remains low, Gemini-3-Pro-Preview-Search reaches only 36.4% accuracy, highlighting persistent failures in cross-image grounding, search triggering, and search strategy coordination. We release the dataset at https: //github.com/fantupang/NAVIGATE.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ff98865a-8f05-4a1e-ad5c-c540baf8c0aa

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines