OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu
摘要
The dream to create AI assistants as capable and versatile as the fictional J.A.R.V.I.S from Iron Man has long captivated imaginations. With the evolution of (multimodal) large language models ((M)LLMs), this dream is closer to reality, as (M)LLM-based Agents using computers, mobile phones and web browsers by operating within the environments and interfaces (e.g., Graphical User Interface (GUI) and Command Line Interface (CLI)) provided by operating systems (OS) to automate tasks have significantly advanced. This paper presents a comprehensive survey on these advanced agents, designated as OS Agents. We begin by elucidating the fundamentals of OS Agents, exploring their key components and capabilities. We then examine methodologies for constructing OS Agents, focusing on domain-specific foundation models and agent frameworks. A detailed review of evaluation metrics and benchmarks highlights how OS Agents are assessed across diverse platforms and tasks. Finally, we discuss current challenges and identify promising directions for future research. An opensource GitHub repository is maintained as a dynamic resource to foster further innovation in this field. † Project Lead, ‡ Core Contributor, * Corresponding Author User: Join the Zoom meeting using the name 'Jack', with ID #303 456 786. OS Agent: Thought: I need to click "Join" button, type the meeting ID and the name "Jack," then click the "Join" button. Action: Click(x=100,y=700), Type('303 456 786'),Type('Jack'), Click(x=100,y=500) 1 2 3 4 Figure 1: An example of OS Agents automatically joining a Zoom meeting on the user's phone as requested.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence On Mobile DevicesDezhi Kong, Zhengzhao Feng, Qiliang Liang, Hao Wang 等CVPR 2026 · 被引用 6 次
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following BehaviorWeikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li 等CVPR 2026 · 被引用 6 次
- Nested Browser-Use Learning for Agentic Information SeekingBaixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li 等ACL 2026 · 被引用 6 次
- LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI AgentsZihe Yan, Zhuosheng Zhang, Jiaping Gui, Gongshen LiuCVPR 2026 · 被引用 5 次
- Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful DemonstrationsZichao Li, Gang Wu, Zichao Wang, Ruiyi Zhang 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper39
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 被引用 539 次
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis 等ICLR 2024 · 被引用 292 次
相关 Paper
- Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHubBohan Lyu, Xin Cong, Heyang Yu, Pan Yang 等ACL 2025
- Agent S: An Open Agentic Framework that Uses Computers Like a HumanSaaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang 等ICLR 2025 · 被引用 2 次
- Spa-Bench: a comprehensive Benchmark for Smartphone Agent EvaluationJingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang 等ICLR 2025
- LLM Agents in Law: Taxonomy, Applications, and ChallengesShuang Liu, Ruijia Zhang, Ruoyun Ma, Yujia Deng 等ACL 2026 · 被引用 3 次
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleRogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont 等ICML 2025
