OSCAR: Operating System Control via State-Aware Reasoning and Re-Planning
Xiaoqiang Wang, Bang Liu
Abstract
Large language models (LLMs) and large multimodal models (LMMs) have shown great potential in automating complex tasks like web browsing and gaming. However, their ability to generalize across diverse applications remains limited, hindering broader utility. To address this challenge, we present OSCAR: Operating System Control via state-Aware reasoning and Re-planning. OSCAR is a generalist agent designed to autonomously navigate and interact with various desktop and mobile applications through standardized controls, such as mouse and keyboard inputs, while processing screen images to fulfill user commands. OSCAR translates human instructions into executable Python code, enabling precise control over graphical user interfaces (GUIs). To enhance stability and adaptability, OSCAR operates as a state machine, equipped with error-handling mechanisms and task-driven re-planning, allowing it to efficiently adjust to real-time feedback and exceptions. We demonstrate OSCAR's effectiveness through extensive experiments on diverse benchmarks across desktop and mobile platforms, where it transforms complex workflows into simple natural language commands, significantly boosting user productivity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45189ef0-84f9-47ae-9372-70103d76a835Cited by top-tier papers4
- GUI-Rise: Structured Reasoning and History Summarization for GUI NavigationTao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu et al.NeurIPS 2025 · 4 citations
- Enhancing GUI Agent with Uncertainty-Aware Self-Trained EvaluatorGongwei Chen, Lirong Jie, Lexiao Zou, Weili Guan et al.NeurIPS 2025 · 4 citations
- The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction GamesZheng Zhang, Deheng Ye, Peilin Zhao, Hao WangACL 2026
- OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser UseXueyu Hu, Tao Xiong, Biao Yi, Zishu Wei et al.ACL 2025
Builds on26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Agent S: An Open Agentic Framework that Uses Computers Like a HumanSaaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang et al.ICLR 2025 · 2 citations
- CoAct-1: Computer-using Multi-agent System with Coding ActionsLinxin Song, Yutong Dai, Viraj Prabhu, Jieyu Zhang et al.ICLR 2026 · 32 citations
- From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use AgentsYuan Wang, Mingyu Li, Haibo ChenEuroSys 2026
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang et al.ICML 2024 · 436 citations
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleRogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont et al.ICML 2025
