ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
Hanyu Lai, Xiao Liu, Yanxiao Zhao, Han Xu, Hanchen Zhang, Bohao Jing, Yanyu Ren, Shuntian Yao, Yuxiao Dong, Jie Tang
Abstract
We introduce COMPUTERRL, a framework for autonomous desktop intelligence that enables agents to operate complex digital workspaces skillfully. COMPUT-ERRL features the API-GUI paradigm, which unifies programmatic API calls and direct GUI interaction to address the inherent mismatch between machine agents and human-centric desktop environments. Scaling end-to-end RL training is crucial for improvement and generalization across diverse desktop tasks; however, it remains challenging due to environmental inefficiency and instability during extended training. To support scalable and robust training, we develop a distributed RL infrastructure capable of orchestrating thousands of parallel virtual desktop environments to accelerate large-scale online RL. Furthermore, we propose Entropulse, a training strategy that alternates reinforcement learning with supervised fine-tuning, effectively mitigating entropy collapse during extended training runs. We employ COMPUTERRL on open models GLM-4-9B-0414 and GLM-4.1V-9B-Thinking, and evaluate them on the OSWorld benchmark. The AUTOGLM-OS-9B achieves a new state-of-the-art accuracy of 48.9%, demonstrating significant improvements for general agents in desktop automation. Our code and the new OFFICEWORLD benchmark are available at our GitHub repository. The algorithm and framework are adopted in building AUTOGLM (Liu et al., 2024b). A u t o G L M -O S 9 B C la u d e 4 .0
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05a87f07-ae8f-4088-95ac-e35ddaeb7fd4Cited by top-tier papers12
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li et al.ICLR 2026 · 54 citations
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI AgentsYifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu et al.ICLR 2026 · 45 citations
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use AgentsHongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu et al.ICLR 2026 · 24 citations
- R-WoM: Retrieval-augmented World Model For Computer-use AgentsKai Mei, Jiang Guo, Shuaichen Chang, Mingwen Dong et al.ICLR 2026 · 12 citations
- ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor CollaborationGaole Dai, Shiqi Jiang, Ting Cao, Yuqing Yang et al.ICLR 2026 · 10 citations
Builds on7
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 484 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- Efficient Agent Training for Computer UseYanheng He, Jiahe Jin, Pengfei LiuICLR 2026 · 15 citations
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu et al.ICLR 2025 · 7 citations
Related papers
- Agent S: An Open Agentic Framework that Uses Computers Like a HumanSaaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang et al.ICLR 2025 · 2 citations
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningZehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai et al.ICLR 2025
- SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand CoresZhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang et al.ICLR 2024 · 10 citations
- ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning EngineeringZexi Liu, Jingyi Chai, Xinyu Zhu, shuo tang et al.ICML 2026
- From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy AssimilationZezhou Wang, Ziyun Zhang, Xiaoyi Zhang, Zhuzhong Qian et al.ACL 2026 · 2 citations
