Efficient Agent Training for Computer Use
Yanheng He, Jiahe Jin, Pengfei Liu
Abstract
Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces reliance on large-scale human demonstrations. Starting with just 312 human-annotated computer use trajectories, we further augment them by synthesizing diverse alternative action decisions with Claude 3.7 Sonnet. Trained on these enriched trajectories, our PC Agent-E model achieved a remarkable 141% relative improvement, and even surpassed the Claude 3.7 Sonnet by 10% in relative terms on WindowsAgentArena-V2, an improved benchmark we also released. By integrating robust human computer use skills with automated AI data synthesis capabilities, our method not only brought substantial improvements over training on human trajectories alone, but also significantly surpassed direct distillation from Claude 3.7 Sonnet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c50cb84-27cc-4f79-8fff-619dcb0dd6f6Cited by top-tier papers4
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li et al.ICLR 2026 · 54 citations
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use AgentsHanyu Lai, Xiao Liu, Yanxiao Zhao, Han Xu et al.ICLR 2026 · 45 citations
- AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State MachinesYifan WU, Yiran Peng, Yiyu Chen, Jianhao Ruan et al.ICML 2026 · 16 citations
- NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI TasksZihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang et al.ACL 2026
Builds on9
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
- Agent S: An Open Agentic Framework that Uses Computers Like a HumanSaaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang et al.ICLR 2025 · 2 citations
- Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionYiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu et al.ICML 2025
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
Related papers
- AgentSynth: Scalable Task Generation for Generalist Computer-Use AgentsJingxu Xie, Dylan Xu, Xuandong Zhao, Dawn SongICLR 2026 · 36 citations
- Watch and Learn: Learning to Use Computers from Online VideosChan Hee Song, Yiwen Song, Palash Goyal, Yu Su et al.CVPR 2026 · 8 citations
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level FilteringYifei He, Pranit Chawla, Yaser Souri, Subhojit Som et al.ACL 2026 · 5 citations
- Anchor: Branch-Point Data Generation for GUI AgentsJinbiao Wei, Yilun Zhao, Kangqi Ni, Arman CohanACL 2026 · 3 citations
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsBofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang et al.AAAI 2026 · 26 citations
