ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, Jianwei Yin
摘要
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding. Recently, with the integration of test-time scaling techniques, these models have also shown strong potential in visual reasoning. However, most existing reasoning approaches remain text-level in nature: MLLMs are prompted to explore various combinations of textual tokens via their underlying language model, while the visual input remains fixed throughout the reasoning process. This paradigm limits the model's ability to fully exploit rich visual information, particularly when dealing with images containing numerous fine-grained elements. In such cases, vision-level reasoning becomes crucial-where models dynamically zoom into specific regions of the image to gather detailed visual cues necessary for accurate decision-making. In this paper, we propose Zoom Eye, a training-free, model-agnostic tree search algorithm tailored for vision-level reasoning. Zoom Eye treats an image as a hierarchical tree structure, where each child node represents a zoomed-in subregion of its parent, and the root corresponds to the full image. The algorithm enables MLLMs to simulate human-like zooming behavior by navigating from root to leaf nodes in search of task-relevant visual evidence. We experiment on a series of elaborate high-resolution benchmarks and the results demonstrate that Zoom Eye not only consistently improves the performance of a series of MLLMs with large margin (e.g., InternVL2.5-8B increases by 15.71% and 17.69% on HR-Bench) but also enables small 3-8B MLLMs to outperform strong large models such as GPT-4o. Our code is available at https://github.com/om-ai-lab/ZoomEye . Zhai et al., 2023) , Multimodal large language models (MLLMs) are able to jointly process textual and visual inputs, achieving impressive performance in vision-language understanding (Zhao et al., 2024a; Bai et al., 2023; Chen et al., 2024b; Li et al., 2024) . Recently, drawing on test-time scaling techniques that enhance reasoning abilities in LLMs, such as OpenAI-o1 (Jaech et al., 2024) and DeepSeek-R1 (Guo et al., 2025), a series of literature tries to investigate these reasoning techniques in MLLMs
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Thyme: Think Beyond ImagesYifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu 等ICLR 2026 · 被引用 146 次
- VGR: Visual Grounded ReasoningJiacong Wang, Zijian Kang, Haochen Wang, Xiao Liang 等ICLR 2026 · 被引用 64 次
- Look-Back: Implicit Visual Re-focusing in MLLM ReasoningShuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye 等AAAI 2026 · 被引用 32 次
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language ReasoningJiaqi Liu, Kaiwen Xiong, Peng Xia, Yiyang Zhou 等ICML 2026 · 被引用 28 次
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong 等ICML 2026 · 被引用 27 次
它引用的顶会 Paper19
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward MechanismCongzhi Zhang, Jiawei Peng, Zhenglin Wang, Yilong Lai 等ACL 2025 · 被引用 6 次
- Pixel Reasoner: Incentivizing Pixel Space Reasoning via Curiosity-Driven Reinforcement LearningAlex Su, Haozhe Wang, Weiming Ren, Fangzhen Lin 等NeurIPS 2025 · 被引用 6 次
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question AnsweringLiangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger 等NeurIPS 2025 · 被引用 23 次
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu 等EMNLP 2025
- Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language ModelsWenshan Wu, Shaoguang Mao, Yadong Zhang, Yan Xia 等NeurIPS 2024 · 被引用 100 次
