LiveNewsBench: Evaluating Web Search Agents with Freshly Curated News
Yunfan Zhang, Kathleen McKeown, Smaranda Muresan
Abstract
Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating such systems remains challenging. We introduce LiveNewsBench, a rigorous and regularly updated benchmark designed to assess the agentic web search abilities of LLMs. LiveNewsBench automatically generates fresh question-answer pairs from recent news articles, ensuring that questions require information beyond an LLM's training data and enabling clear separation between internal knowledge and search capability. The benchmark features intentionally difficult questions requiring multi-hop search queries, page visits, and reasoning, making it well-suited for evaluating agentic search behavior. Our automated data curation and question generation pipeline enables frequent benchmark updates and supports construction of a large-scale training dataset for agentic web search models, addressing the scarcity of such data in the research community. To ensure reliable evaluation, we include a subset of human-verified samples in the test set. We evaluate a broad range of systems using LiveNewsBench, including commercial and open-weight LLMs as well as LLM-based web search APIs. The leaderboard, datasets, and code are publicly available at livenewsbench.com.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a3598d6-7b3b-437d-b2a0-d869d5171633Builds on5
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- LiveBench: A Challenging, Contamination-Limited LLM BenchmarkColin White, Samuel Dooley, Manley Roberts, Arka Pal et al.ICLR 2025
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li et al.ICLR 2025
Related papers
- WideSearch: Benchmarking Agentic Broad Info-SeekingRyan Wong, Jiawei Wang, Junjie Zhao, Li Chen et al.ICLR 2026 · 66 citations
- AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic WebShanshan Zhong, Kate Shen, Chenyan XiongICML 2026 · 2 citations
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?Zhanke Zhou, Xiao Feng, Zhaocheng Zhu, Jiangchao Yao et al.ICML 2025
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin et al.EMNLP 2024 · 5 citations
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
