BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents
Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Sahel Sharifymoghaddam, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, Yanxi Li
Abstract
Deep search agents that combine large language models with retrieval tools excel at complex, multi-hop queries. Yet, existing benchmarks such as BrowseComp rely on black-box web search APIs, facing key limitations. (1) Fairness : for agents, dynamic and opaque web APIs hinder reproducibility and fair comparisons across agents. (2) Disentanglement : for retrieval, the lack of a fixed document corpus makes it impossible to isolate retriever contributions from end-to-end search agent accuracy. We introduce BrowseComp-Plus , a benchmark derived from BrowseComp that employs a fixed, human-verified corpus, enabling controlled retrieval for deep search agents. BrowseComp-Plus clearly distinguishes agent performance: with a BM25 retriever, the open-source Search-R1 achieves 3.86% accuracy, while GPT-5 achieves 55.9%. Additionally, BrowseComp-Plus makes retrieval gains explicit: pairing GPT-5 with Qwen3-Embedding-8B retriever further improves accuracy to 70.1% while reducing search calls. Overall, BrowseComp-Plus provides a fair and disen-tangled testbed, advancing both deep search agent evaluation and retrieval research for agen-tic search. Code and data can be found at: https://texttron.github.io/BrowseComp-Plus/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- RA-DIT: Retrieval-Augmented Dual Instruction TuningXi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi et al.ICLR 2024 · 229 citations
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 211 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
Related papers
- InteractComp: Evaluating Search Agents With Ambiguous QueriesMingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong et al.ICML 2026 · 11 citations
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentXinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang et al.ICLR 2026 · 79 citations
- DRBench: A Realistic Benchmark for Enterprise Deep ResearchAmirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox et al.ICLR 2026 · 18 citations
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu et al.ICLR 2026 · 29 citations
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
