DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
Shengqin Wang, Wentao Yan, Huichi Zhou, Yihang Chen, Kun Shao, Zhizhong Zhang, Yuan Xie
Abstract
Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is observed that such agents often meet premature interaction collapse, caused by two primary reasons: 1) the terminal reward often appending on the last token prevents the advantage from distinguishing trajectories with exploratory behavior; 2) excessively redundant context hinders the agent from absorbing useful feedback. To address these issues, we propose the Deepening Reasoning MMSearchAgent, the framework leverages the structural proximity to derive advantage signals from the whole rollout trajectories in an entire batch, such that trajectories of different lengths are further encouraged to be generated, even when containing the same correct answer. Additionally, differentiated gaussian rewards are employed to dynamically calibrate interaction tolerance, thereby ensuring information reliability and reducing redundancy. To support multi-turn interaction training, we have constructed a multi-step deep-reasoning dataset including 3602 high-quality QA pairs with at least 3 reasoning steps. Extensive experiments demonstrate that our method achieves state-of-the-art performance, outperforming the MMSearch-R1 by 8.4% on FVQA-test.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccf15a91-9ac9-4e35-892b-da041ffc9bc6Builds on13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu et al.NeurIPS 2025 · 356 citations
- RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsYue Yu, Wei Ping, Zihan Liu, Boxin Wang et al.NeurIPS 2024 · 321 citations
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-ImprovementXiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu et al.NeurIPS 2025 · 158 citations
- MMSearch-R1: Incentivizing LMMs to SearchJinming Wu, Zihao Deng, Wei Li, Yiding Liu et al.ACL 2026 · 93 citations
Related papers
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMsShreyas Singh, Kunal Singh, Pradeep MoturiICLR 2026 · 5 citations
- MM-DeepResearch: A Simple and Effective Multimodal Agentic Search BaselineHuanjin Yao, Qixiang Yin, Min Yang, Ziwang Zhao et al.ICML 2026 · 14 citations
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsJunde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu et al.ACL 2025 · 88 citations
- Hybrid Deep Searcher: Scalable Parallel and Sequential Search ReasoningDayoon Ko, Jihyuk Kim, Haeju Park, Sohyeon Kim et al.ICLR 2026 · 2 citations
- Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner CollaborationBowei He, Minda Hu, Zenan Xu, Hongru WANG et al.ICML 2026 · 9 citations
