AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, Jonathan Berant
Abstract
Language agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web. In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets or locating relevant nearby businesses. We introduce ASSISTANTBENCH, a challenging new benchmark consisting of 214 realistic tasks that can be automatically evaluated, covering different scenarios and domains. We find that AS-SISTANTBENCH exposes the limitations of current systems, including language models and retrieval-augmented language models, as no model reaches an accuracy of more than 26 points. While closed-book LMs perform well in terms of accuracy, they exhibit low precision and tend to hallucinate facts. State-of-the-art web agents reach a score of near zero. Additionally, we introduce SEEPLANACT (SPA), a new web agent that significantly outperforms previous agents, and an ensemble of SPA and closed-book models reaches the best overall performance. Moreover, we analyze failures of current systems and highlight that open web navigation remains a major challenge. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir et al.ICLR 2026 · 86 citations
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web AgentsIdo Levy, Ben wiesel, Sami Marreed, Alon Oved et al.ICLR 2026 · 78 citations
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent SystemsMing Ma, Jue Zhang, Fangkai Yang, Yu Kang et al.ICLR 2026 · 24 citations
- OpenApps: Simulating Environment Variations to Measure UI Agent ReliabilityKaren Ullrich, Jingtong Su, Claudia Shi, Arjun Subramonian et al.ICLR 2026 · 10 citations
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li et al.ACL 2026 · 8 citations
Builds on27
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
- LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context GrowthWeihao Zeng, Yuzhen Huang, Junxian HeICML 2026 · 11 citations
- GTA: Generating Long-horizon Tasks for Web Agents at ScaleTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou et al.ACL 2026
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
