CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
Liupeng Li, Haoqian Kang, Zhenyu Lu, Jinpeng Wang, Bin Chen, Ke Chen, Yaowei Wang
Abstract
High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind spots when proposals fail, whereas scan-based search guarantees coverage at the cost of computational redundancy and semantic fragmentation. To address this dilemma, we introduce CVSearch, a training-free adaptive framework that dynamically schedules search strategies via an Assess-then-Search workflow. Specifically, CVSearch first invokes expertassisted search when global information is insufficient, and only triggers a novel semanticaware scanning mechanism upon failure. Distinct from rigid grid partitioning, this efficient scanning paradigm incorporates Semantic Guided Adaptive Patching to decompose images into semantically consistent regions, effectively mitigating object fragmentation. Furthermore, we devise a Dynamic Bottom-Up Search strategy driven by a Visual Complexity prior to enable efficient and precise iterative exploration of local details. Extensive experiments on HR benchmarks demonstrate that CVSearch achieves state-of-the-art accuracy while substantially improving search efficiency. Code is released at ICML26-CVSearch. grained details essential for real-world tasks (Zhang et al., 2024a), as illustrated in Figure 1 (a). To mitigate this limitation, recent research has branched into three paradigms: (1) Cropping-based paradigms (Li
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 997f88bd-e8bf-44e6-96fa-2ae613b282daBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma et al.CVPR 2024 · 111 citations
- Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language ModelsWenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou et al.AAAI 2025 · 4 citations
Related papers
- BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained PerceptionGeng Li, Yuxin PengICML 2026
- Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAGWenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang et al.ICML 2025
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical DecouplingXianjie Liu, Yiman Hu, Yixiong Zou, Liang Wu et al.ICML 2026 · 14 citations
- ActiveScope: Actively Seeking and Correcting Perception for MLLMsYajing Wang, Chao Bi, Junshu Sun, Shufan Shen et al.ICML 2026 · 1 citation
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question AnsweringLiangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger et al.NeurIPS 2025 · 23 citations
