Do LLM Agents Know How to Ground, Recover, and Assess? Evaluating Epistemic Competence in Information-Seeking Agents
Jiaqi Shao, Yuxiang Lin, Munish Prasad Lohani, Yufeng Miao, Bing Luo
摘要
Recent work has explored training Large Language Model (LLM) search agents with reinforcement learning (RL) for open-domain question answering. However, most evaluations focus solely on final answer accuracy, overlooking how these agents reason with and act on external evidence. We introduce SeekBench, the first process-level evaluation framework for LLM search agents that operationalize epistemic competence through metrics derived from an annotation schema. We develop and validate our annotation schema using an expert-annotated dataset of 190 traces (over 1,800 steps). To evaluate at scale, we introduce an LLM-as-judge pipeline. Our framework provides granular analysis of whether agents demonstrate: (1) groundedness, by generating reasoning steps supported by observed evidence; (2) recovery, by adaptively reformulating searches to recover from low-quality results; and (3) calibration, by correctly assessing whether current evidence is sufficient to provide an answer. By applying our evaluation framework to state-of-the-art search agents tuned on Qwen2.5-7B, we uncover critical behavioral gaps that answer-only metrics miss, as well as specialized skills such as Search-R1's synthesis abilities. These analyses highlight distinct epistemic competencies, offering actionable insights for the development of more capable and trustworthy agents. Code is available at https://github.com/SHAO-Jiaqi757/SeekBench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
- Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive SearchMaohao Shen, Guangtao Zeng, Zhenting Qi, Zhang-Wei Hong 等ICML 2025
相关 Paper
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 被引用 1 次
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun 等NeurIPS 2025 · 被引用 125 次
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive RewardsShiyu Li, Yifan Wang, Peiming Li, Zheng Wei 等ICML 2026 · 被引用 8 次
- Reflection-Bench: Evaluating Epistemic Agency in Large Language ModelsLingyu Li, Yixu Wang, Haiquan Zhao, Shuqi Kong 等ICML 2025
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 被引用 2 次
