WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, Alexandre Lacoste
摘要
We study the use of large language model-based agents for interacting with software via web browsers. Unlike prior work, we focus on measuring the agents' ability to perform tasks that span the typical daily work of knowledge workers utilizing enterprise software systems. To this end, we propose WorkArena, a remote-hosted benchmark of 33 tasks based on the widely-used ServiceNow platform. We also introduce BrowserGym, an environment for the design and evaluation of such agents, offering a rich set of actions as well as multimodal observations. Our empirical evaluation reveals that while current agents show promise on WorkArena, there remains a considerable gap towards achieving full task automation. Notably, our analysis uncovers a significant performance disparity between open and closed-source LLMs, highlighting a critical area for future exploration and development in the field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at ScaleTianyue Ou, Frank F. Xu, Aman Madaan, Jiarui Liu 等NeurIPS 2024 · 被引用 45 次
- Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsHyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim 等NeurIPS 2025 · 被引用 37 次
- macOSWorld: A Multilingual Interactive Benchmark for GUI AgentsPei Yang, Hai Ci, Mike Zheng ShouNeurIPS 2025 · 被引用 34 次
- How to Train Your LLM Web Agent: A Statistical DiagnosisDheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei 等NeurIPS 2025 · 被引用 19 次
- ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT PipelinesTengjun Jin, Yuxuan Zhu, Daniel KangVLDB 2026 · 被引用 13 次
它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 被引用 539 次
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo 等ICLR 2024 · 被引用 160 次
- A data-driven approach for learning to control computersPeter Conway Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton 等ICML 2022 · 被引用 124 次
相关 Paper
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleRogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont 等ICML 2025
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur 等ACL 2024 · 被引用 25 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong 等ACL 2025 · 被引用 20 次
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
