CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
Liupeng Li, Haoqian Kang, Zhenyu Lu, Jinpeng Wang, Bin Chen, Ke Chen, Yaowei Wang
摘要
High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind spots when proposals fail, whereas scan-based search guarantees coverage at the cost of computational redundancy and semantic fragmentation. To address this dilemma, we introduce CVSearch, a training-free adaptive framework that dynamically schedules search strategies via an Assess-then-Search workflow. Specifically, CVSearch first invokes expertassisted search when global information is insufficient, and only triggers a novel semanticaware scanning mechanism upon failure. Distinct from rigid grid partitioning, this efficient scanning paradigm incorporates Semantic Guided Adaptive Patching to decompose images into semantically consistent regions, effectively mitigating object fragmentation. Furthermore, we devise a Dynamic Bottom-Up Search strategy driven by a Visual Complexity prior to enable efficient and precise iterative exploration of local details. Extensive experiments on HR benchmarks demonstrate that CVSearch achieves state-of-the-art accuracy while substantially improving search efficiency. Code is released at ICML26-CVSearch. grained details essential for real-world tasks (Zhang et al., 2024a), as illustrated in Figure 1 (a). To mitigate this limitation, recent research has branched into three paradigms: (1) Cropping-based paradigms (Li
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
- Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language ModelsWenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou 等AAAI 2025 · 被引用 4 次
相关 Paper
- BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained PerceptionGeng Li, Yuxin PengICML 2026
- Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAGWenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang 等ICML 2025
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical DecouplingXianjie Liu, Yiman Hu, Yixiong Zou, Liang Wu 等ICML 2026 · 被引用 14 次
- ActiveScope: Actively Seeking and Correcting Perception for MLLMsYajing Wang, Chao Bi, Junshu Sun, Shufan Shen 等ICML 2026 · 被引用 1 次
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question AnsweringLiangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger 等NeurIPS 2025 · 被引用 23 次
