AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, Huzefa Rangwala
摘要
Autonomy via agents using large language models (LLMs) for personalized, standardized tasks boosts human efficiency. Automating web tasks (like booking hotels within a budget) is increasingly sought after. Fulfilling practical needs, the web agent also serves as an important proof-of-concept example for various agent grounding scenarios, with its success promising advancements in many future applications. Prior research often handcrafts web agent strategies (e.g., prompting templates, multi-agent systems, search methods, etc.) and the corresponding in-context examples, which may not generalize well across all real-world scenarios. On the other hand, there has been limited study on the misalignment between a web agent's observation/action representation and the pre-training data of the LLM it's based on. This discrepancy is especially notable when LLMs are primarily trained for language completion rather than tasks involving embodied navigation actions and symbolic web elements. Our study enhances an LLM-based web agent by simply refining its observation and action space to better align with the LLM's capabilities. This approach enables our base agent to significantly outperform previous methods on a wide variety of web tasks. Specifically, on WebArena, a benchmark featuring general-purpose web interaction tasks, our agent AgentOccam surpasses the previous state-of-the-art and concurrent work by 9.8 (+29.4%) and 5.9 (+15.8%) absolute points respectively, and boosts the success rate by 26.6 points (+161%) over similar plain web agents with its observation and action space alignment. We achieve this without using in-context examples, new agent roles, online feedback or search strategies. AgentOccam's simple design highlights LLMs' impressive zero-shot performance on web tasks, and underlines the critical role of carefully tuning observation and action spaces for LLM-based agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web AgentsIdo Levy, Ben wiesel, Sami Marreed, Alon Oved 等ICLR 2026 · 被引用 78 次
- Thinking vs. Doing: Improving Agent Reasoning by Scaling Test-Time InteractionJunhong Shen, Hao Bai, Lunjun Zhang, Yifei Zhou 等NeurIPS 2025 · 被引用 34 次
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making TasksVishnu Sarukkai, Zhiqiang Xie, Kayvon FatahalianNeurIPS 2025 · 被引用 22 次
- PlugMem: A Task-Agnostic Plugin Memory Module for LLM AgentsKe Yang, Zixi Chen, Xuan He, Jize Jiang 等ICML 2026 · 被引用 20 次
- Test-Time Adaptation for LLM Agents via Environment InteractionArthur Chen, Zuxin Liu, Jianguo Zhang, Akshara Prabhakar 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper11
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang 等ICML 2024 · 被引用 443 次
- AdaPlanner: Adaptive Planning from Feedback with Language ModelsHaotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai 等NeurIPS 2023 · 被引用 257 次
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer ControlLongtao Zheng, Rundong Wang, Xinrun Wang, Bo AnICLR 2024 · 被引用 132 次
相关 Paper
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur 等ACL 2024 · 被引用 25 次
- Go-Browse: Training Web Agents with Structured ExplorationApurva Gandhi, Graham NeubigICLR 2026 · 被引用 30 次
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web NavigationHyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak 等ICLR 2025 · 被引用 2 次
- Learning to Contextualize Web Pages for Enhanced Decision Making by LLM AgentsDongjun Lee, Juyong Lee, Kyuyoung Kim, Jihoon Tack 等ICLR 2025
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
