InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
Kaican Li, Lewei Yao, Jiannan Wu, Tiezheng YU, Jierun Chen, Haoli Bai, Lu Hou, Lanqing HONG, Wei Zhang, Nevin L. Zhang
摘要
The ability for AI agents to "think with images" requires a sophisticated blend of reasoning and perception. However, current open multimodal agents still largely fall short on the reasoning aspect crucial for real-world tasks like analyzing documents with dense charts/diagrams and navigating maps. To address this gap, we introduce O3-bench, a new benchmark designed to evaluate multimodal reasoning with interleaved attention to visual details. O3-bench features challenging problems that require agents to piece together subtle visual information from distinct image areas through multi-step reasoning. The problems are highly challenging even for frontier systems like OpenAI o3, which only obtains 40.8% accuracy on O3-bench. To make progress, we propose InSight-o3, a multi-agent framework consisting of a visual reasoning agent (vReasoner) and a visual search agent (vSearcher) for which we introduce the task of generalized visual search---locating relational, fuzzy, or conceptual regions described in free-form language, beyond just simple objects or figures in natural images. We then present a multimodal LLM purpose-trained for this task via reinforcement learning. As a plug-and-play agent, our vSearcher empowers frontier multimodal models (as vReasoners), significantly improving their performance on a wide range of benchmarks. This marks a concrete step towards powerful o3-like open systems. Our code and dataset can be found at https://github.com/m-Just/InSight-o3.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
相关 Paper
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language ModelsYuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang 等CVPR 2025
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and MethodHaochen Wang, Xiangtai Li, Zilong Huang, Anran Wang 等ICLR 2026 · 被引用 45 次
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual SearchXin Lai, Junyi Li, Wei Li, Tao Liu 等ICLR 2026 · 被引用 124 次
- GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical EvaluationYuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao 等ICLR 2026 · 被引用 4 次
- MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning ModelsVanya Cohen, Ray MooneyICML 2026 · 被引用 2 次
