MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
Huanjin Yao, Qixiang Yin, Min Yang, Ziwang Zhao, Yibo Wang, Haotian Luo, Jingyi Zhang, Jiaxing Huang
Abstract
We aim to develop a multimodal research agent capable of explicit reasoning and planning, multitool invocation, and cross-modal information synthesis, enabling it to conduct deep research tasks. However, we observe three main challenges in developing such agents: (1) scarcity of searchintensive multimodal QA data, (2) lack of effective search trajectories, and (3) prohibitive cost of training with online search APIs. To tackle them, we first propose Hyper-Search, a hypergraphbased QA generation method that models and connects visual and textual nodes within and across modalities, enabling to generate search-intensive multimodal QA pairs that require invoking various search tools to solve. Second, we introduce DR-TTS, which first decomposes searchinvolved tasks into several categories according to search tool types, and respectively optimize specialized search tool experts for each tool. It then recomposes tool experts to jointly explore search trajectories via tree search, producing trajectories that successfully solve complex tasks using various search tools. Third, we build an offline search engine supporting multiple search tools, enabling agentic reinforcement learning without using costly online search APIs. With the three designs, we develop MM-DeepResearch, a powerful multimodal deep research agent, and extensive results shows its superiority across benchmarks. Code is available at https://github.com/ HJYao00/MM-DeepResearch
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 034ea388-73b4-46b7-853c-37edc5f0e7ceCited by top-tier papers1
Ask how each one uses itBuilds on15
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- MMSearch-R1: Incentivizing LMMs to SearchJinming Wu, Zihao Deng, Wei Li, Yiding Liu et al.ACL 2026 · 93 citations
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsJunde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu et al.ACL 2025 · 88 citations
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningQiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen et al.NeurIPS 2025 · 76 citations
- Dense Connector for MLLMsHuanjin Yao, Wenhao Wu, Taojiannan Yang, Yuxin Song et al.NeurIPS 2024 · 65 citations
Related papers
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentXinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang et al.ICLR 2026 · 79 citations
- Unlocking Long-Horizon Agentic Search with Large-Scale End-to-End RLJiaxuan Gao, Wei Fu, Minyang Xie, Shusheng Xu et al.ICLR 2026
- DeepEyesV2: Toward Agentic Multimodal ModelJack Hong, Chenxiao Zhao, ChengLIn Zhu, Weiheng Lu et al.ICLR 2026 · 109 citations
- DR-MMSearchAgent: Deepening Reasoning in Multimodal Search AgentsShengqin Wang, Wentao Yan, Huichi Zhou, Yihang Chen et al.ICML 2026
- IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement LearningHaohao Luo, Zexi Li, Yuexiang Xie, Wenhao Zhang et al.ICML 2026
