Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, Caiming Xiong
Abstract
Automating GUI tasks remains challenging due to reliance on textual representations, platformspecific action spaces, and limited reasoning capabilities. We introduce AGUVIS, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct AGUVIS DATA COLLECTION, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that AGUVIS achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project. github.io to advance future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23997e26-7e60-4c0d-95ea-a6c1b1710f18Cited by top-tier papers88
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang et al.NeurIPS 2025 · 151 citations
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang et al.ICLR 2026 · 109 citations
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin et al.AAAI 2026 · 103 citations
- GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsQianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang et al.NeurIPS 2025 · 98 citations
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou et al.NeurIPS 2025 · 73 citations
Builds on15
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 539 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari et al.ICLR 2024 · 359 citations
Related papers
- UIPro: Unleashing Superior Interaction Capability for GUI AgentsHongxin Li, Jingran Su, Jingfan Chen, Zheng Ju et al.ICCV 2025
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- OS-ATLAS: Foundation Action Model for Generalist GUI AgentsZhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang et al.ICLR 2025
- Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI AgentsBoyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie et al.ICLR 2025
- GUICourse: From General Vision Language Model to Versatile GUI AgentWentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin et al.ACL 2025
