InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
Kaican Li, Lewei Yao, Jiannan Wu, Tiezheng YU, Jierun Chen, Haoli Bai, Lu Hou, Lanqing HONG, Wei Zhang, Nevin L. Zhang
Abstract
The ability for AI agents to "think with images" requires a sophisticated blend of reasoning and perception. However, current open multimodal agents still largely fall short on the reasoning aspect crucial for real-world tasks like analyzing documents with dense charts/diagrams and navigating maps. To address this gap, we introduce O3-bench, a new benchmark designed to evaluate multimodal reasoning with interleaved attention to visual details. O3-bench features challenging problems that require agents to piece together subtle visual information from distinct image areas through multi-step reasoning. The problems are highly challenging even for frontier systems like OpenAI o3, which only obtains 40.8% accuracy on O3-bench. To make progress, we propose InSight-o3, a multi-agent framework consisting of a visual reasoning agent (vReasoner) and a visual search agent (vSearcher) for which we introduce the task of generalized visual search---locating relational, fuzzy, or conceptual regions described in free-form language, beyond just simple objects or figures in natural images. We then present a multimodal LLM purpose-trained for this task via reinforcement learning. As a plug-and-play agent, our vSearcher empowers frontier multimodal models (as vReasoners), significantly improving their performance on a wide range of benchmarks. This marks a concrete step towards powerful o3-like open systems. Our code and dataset can be found at https://github.com/m-Just/InSight-o3.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da2070a8-e0fd-4a4e-9a11-31b5e8e879f5Builds on30
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
Related papers
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language ModelsYuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang et al.CVPR 2025
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and MethodHaochen Wang, Xiangtai Li, Zilong Huang, Anran Wang et al.ICLR 2026 · 45 citations
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual SearchXin Lai, Junyi Li, Wei Li, Tao Liu et al.ICLR 2026 · 124 citations
- GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical EvaluationYuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao et al.ICLR 2026 · 4 citations
- MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning ModelsVanya Cohen, Ray MooneyICML 2026 · 2 citations
