MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanmin Wu, Jiayi Lei, Pengshuo Qiu, Pan Lu, Zehui Chen, Guanglu Song, Peng Gao, Yu Liu, Chunyuan Li, Hongsheng Li
摘要
The advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleaved nature of website information. Recently, Large Multimodal Models (LMMs) have made impressive strides. Yet, whether they can function as AI search engines remains under-explored, leaving the potential of LMMs in multimodal search an open question. To this end, we first design a delicate pipeline, MMSearch-Engine, to empower any LMMs with multimodal search capabilities. On top of this, we introduce MMSearch, a comprehensive evaluation benchmark to assess the multimodal search performance of LMMs. The curated dataset contains 300 manually collected instances spanning 14 subfields, which involves no overlap with the current LMMs' training data, ensuring the correct answer can only be obtained within searching. By using MMSearch-Engine, the LMMs are evaluated by performing three individual tasks (requery, rerank, and summarization), and one challenging end-to-end task with a complete searching process. We conduct extensive experiments on closed-source and open-source LMMs. Among all tested models, GPT-4o with MMSearch-Engine achieves the best results, which surpasses the commercial product, Perplexity Pro, in the end-to-end task, demonstrating the effectiveness of our proposed pipeline. We further present error analysis to unveil current LMMs still struggle to fully grasp the multimodal search tasks, and conduct ablation study to indicate the potential of scaling test-time computation for AI search engine. We hope MMSearch may provide unique insights to guide the future development of multimodal AI search engine.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language ModelsWenxuan Huang, Yu Zeng, Qiuchen Wang, Zhen Fang 等ICML 2026 · 被引用 27 次
- Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?Naen Xu, Jinghuai Zhang, Changjiang Li, Hengyu An 等AAAI 2026 · 被引用 6 次
- Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual UnderstandingJunhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang 等CVPR 2026 · 被引用 2 次
- Answering Narrative-Driven Recommendation Queries via a Retrieve-Rank Paradigm and the OCG-AgentYunxiao Shi, Haoning Shang, Xing Zi, Wujiang Xu 等EMNLP 2025 · 被引用 1 次
- Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based DocumentsYuming Yang, Jiang Zhong, Li Jin, Xiao Sun 等ACL 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- MindSearch: Mimicking Human Minds Elicits Deep AI SearcherZehui Chen, Kuikun Liu, Qiuchen Wang, Jiangning Liu 等ICLR 2025 · 被引用 2 次
- MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language ModelsFanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu 等ICLR 2025
- INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in InsuranceChenwei Lin, Hanjia Lyu, Xian Xu, Jiebo LuoICCV 2025 · 被引用 3 次
- MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingFei Wang, Xingyu Fu, James Y. Huang, Zekun Li 等ICLR 2025
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao 等ICML 2025
