Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
Weihao Zeng, Keqing He, Chuqiao Kuang, Xiaoguang Li, Junxian He
摘要
Test-time compute can be scaled both sequentially and in parallel. Sequential scaling involves lengthening the generation process, while parallel scaling involves verifying and selecting among multiple candidate outputs. Combining these two strategies has led to the most powerful AI systems, such as Grok 4 Heavy, GPT-5 Pro, and Gemini-2.5 Pro Deep Think. A key observation is that, in certain contexts (e.g., solving Sudoku puzzles), verifying responses can be substantially easier than generating them. This property, referred to as asymmetric verification, highlights the strong potential of test-time scaling. In this work, we study both sequential and parallel test-time scaling of deep search agents, motivated by the intuition that verification in this setting is often much easier than generation. In experiments, we first show that sequential scaling methods, such as budget forcing, can be effective initially but eventually degrade performance when over-applied in agentic search. Due to asymmetric verification, however, we are able to achieve substantial improvements by allocating only a modest amount of compute to the verifier. We conduct experiments with flagship open-source models, including GLM-4.5, K2, Qwen3-2507 and Tongyi-DeepResearch, and extend them to their ``Heavy'' variants through test-time scaling. These deep research agents achieve improvements of up to 20 absolute points on benchmarks such as BrowseComp. Remarkably, as an open-source alternative, GLM-4.5 Heavy reaches accuracy of 54.0% on BrowseComp, 66.0% on GAIA, and 68.0% on xbench-DeepSearch, placing it on par with the best proprietary choices such as OpenAI Deep Research and o3. Tongyi-DeepResearch Heavy pushes performance even further, attaining 69.0% accuracy on BrowseComp.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InteractComp: Evaluating Search Agents With Ambiguous QueriesMingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong 等ICML 2026 · 被引用 11 次
- Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video UnderstandingHaiyang Yan, Hongyun Zhou, Peng Xu, Xiaoxue Feng 等CVPR 2026 · 被引用 8 次
- Deep Search with Hierarchical Meta-Cognitive Monitoring Inspired by Cognitive NeuroscienceZhongxiang Sun, Qipeng Wang, Weijie Yu, Jingxuan Yang 等SIGIR 2026
它引用的顶会 Paper5
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
- B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught ReasonersWeihao Zeng, Yuzhen Huang, Lulu Zhao, Yijun Wang 等ICLR 2025
- Diving into Self-Evolving Training for Multimodal ReasoningWei Liu, Junlong Li, Xiwen Zhang, Fan Zhou 等ICML 2025
相关 Paper
- Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling VerificationEric Zhao, Pranjal Awasthi, Sreenivas GollapudiICML 2025
- Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time ScalingHao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo 等NeurIPS 2025 · 被引用 8 次
- EGSS: Entropy-guided Stepwise Scaling for Reliable Software EngineeringChenhui Mao, Yuanting Lei, Zhixiang Wei, Ming Liang 等ACL 2026
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen 等ACL 2026 · 被引用 3 次
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning ModelsSoumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu 等NeurIPS 2025 · 被引用 43 次
