ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, Mike Zheng Shou
Abstract
Building Graphical User Interface (GUI) assistants holds significant promise for enhancing human workflow productivity. While most agents are language-based, relying on closed-source API with text-rich meta-information (e.g., HTML or accessibility tree), they show limitations in perceiving UI visuals as humans do, highlighting the need for GUI visual agents. In this work, we develop a visionlanguage-action model in digital world, namely ShowUI, which features the following innovations: (i) UI-Guided Visual Token Selection to reduce computational costs by formulating screenshots as an UI connected graph, adaptively identifying their redundant relationship and serve as the criteria for token selection during self-attention blocks; (ii) Interleaved Vision-Language-Action Streaming that flexibly unifies diverse needs within GUI tasks, enabling effective management of visual-action history in navigation or pairing multi-turn query-action sequences per screenshot to enhance training efficiency; (iii) Small-scale Highquality GUI Instruction-following Datasets by careful data curation and employing a resampling strategy to address significant data type imbalances. With above components, ShowUI, a lightweight 2B model using 256K data, achieves a strong 75.1% accuracy in zero-shot screenshot grounding. Its UI-guided token selection further reduces 33% of redundant visual tokens during training and speeds up the performance by 1.4×. Navigation experiments across web [12], mobile [35], and online [39] environments further underscore the effectiveness and potential of our model in advancing GUI visual agents. The models are available at https://github.com/showlab/ShowUI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers51
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang et al.ICLR 2026 · 109 citations
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li et al.ICLR 2026 · 54 citations
- GUI-G²: Gaussian Reward Modeling for GUI GroundingFei Tang, Zhangxuan Gu, Zhengxi Lu, Xuyang Liu et al.AAAI 2026 · 48 citations
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsQiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding et al.ICLR 2026 · 45 citations
- MobileUse: A Hierarchical Reflection-Driven GUI Agent for Autonomous Mobile OperationNing Li, Xiangmou Qu, Jiamu Zhou, Muning Wen et al.NeurIPS 2025 · 42 citations
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari et al.ICLR 2024 · 359 citations
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionRui Yang, Lin Song, Yanwei Li, Sijie Zhao et al.NeurIPS 2023 · 340 citations
Related papers
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- FocusUI: Efficient UI Grounding via Position-Preserving Visual Token SelectionMingyu Ouyang, Kevin Qinghong Lin, Mike Zheng Shou, Hwee Tou NgCVPR 2026 · 8 citations
- SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language ModelZixin Guo, Yue Jiang, Luis A. Leiva, Antti OulasvirtaCHI 2026 · 1 citation
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI AgentsBoyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie et al.ICLR 2025
