A Survey of Large Language Model-Based Search Agents
Yunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou, Rong Shan, Te Gao, Jiachen Zhu, Weiwen Liu, Yong Yu, Weinan Zhang
摘要
The advent of Large Language Models (LLMs) has significantly revolutionized web search. The emergence of LLM-based Search Agents marks a pivotal shift towards deeper, dynamic, autonomous information seeking. These agents can comprehend user intentions and environmental context and execute multi-turn retrieval with dynamic planning, extending search capabilities far beyond the web. Leading examples like OpenAI's Deep Research highlight their potential for deep information mining and realworld applications. This survey provides the first systematic analysis of search agents. We comprehensively analyze and categorize existing works from the perspectives of architecture, optimization, application, and evaluation, ultimately identifying critical open challenges and outlining promising future research directions in this rapidly evolving field. Our repository is available on https://github.com/ YunjiaXi/Awesome-Search-Agent-Papers . * Corresponding author 1 8244 sub-domain or perspective, e.g., Deep Research which emphasizes professional report generation from extensive information seeking (Xu and Peng, 2025; Huang et al., 2025b) or the integration of reasoning and RAG (Liang et al., 2025; Gao et al., 2025) , our work comprehensively analyzes the holistic pipeline of search agents, including search structure, optimization, application, evaluation, and challenges. For each part, we provide a thorough analysis of representative works and developing tendencies. Specifically, this paper is structured as follows: Sec. 2 introduces the task formulation for search agents. How to Search in Sec. 3 presents how agents scale up search turns and utilize complex search structures (i.e., parallel, sequential, and hybrid) to determine query content. How to Optimize in Sec. 4 discusses various optimization methodologies for search agents, including tuning and nontuning approaches. How to Apply in Sec. 5 delineates the extensive application areas of search agents, encompassing both internal agent enhancements (e.g., reasoning, memory, and tool-use) and external applications (e.g., math, medicine, and finance). How to Evaluate in Sec. 5 introduces evaluation of search agents, covering various datasets and metrics. Finally, Sec. 7 presents current challenges and promising future research directions. Task Formulation Given a user's intention q and context C, a search agent iteratively plans and acts to gather information and fulfill the user's intention. Upon receiving intention q, the agent initiates a planning π 0 = Plan(q, C) to conduct an information seeking trajectory. At each step t, the agent reflects on its current observation o t and previous trajectory t and updates its plan π t+1 = Reflect(o t , h t , π t ). It then performs an action a t+1 = Act(π t+1 ) (e.g., search for or browse certain content) yielding a new observation o t , e.g., the retrieved information. This process continues until sufficient information is acquired, forming a sequence of observations O = o 1 , o 2 , • • • , o T . From O, the agent extracts and ranks the most relevant data into an evidence set E = Select(q, O) and generates a response ŷq = Generate(q, E) to fulfill the user's intention.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian 等NeurIPS 2025 · 被引用 354 次
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun 等EMNLP 2023 · 被引用 315 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
相关 Paper
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsYuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai 等EMNLP 2025 · 被引用 8 次
- Learning to Retrieve from Agent TrajectoriesYuqi Zhou, Sunhao Dai, Changle Qu, Liang Pang 等SIGIR 2026
- WideSearch: Benchmarking Agentic Broad Info-SeekingRyan Wong, Jiawei Wang, Junjie Zhao, Li Chen 等ICLR 2026 · 被引用 66 次
- LLM Agents in Law: Taxonomy, Applications, and ChallengesShuang Liu, Ruijia Zhang, Ruoyun Ma, Yujia Deng 等ACL 2026 · 被引用 3 次
- MindSearch: Mimicking Human Minds Elicits Deep AI SearcherZehui Chen, Kuikun Liu, Qiuchen Wang, Jiangning Liu 等ICLR 2025 · 被引用 2 次
