FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
Mingyu Ouyang, Kevin Qinghong Lin, Mike Zheng Shou, Hwee Tou Ng
Abstract
Vision-Language Models (VLMs) have shown remarkable performance in User Interface (UI) grounding tasks, driven by their ability to process increasingly high-resolution screenshots. However, screenshots are tokenized into thousands of visual tokens (e.g., about 4700 for 2K resolution), incurring significant computational overhead and diluting attention. In contrast, humans typically focus on regions of interest when interacting with UI. In this work, we pioneer the task of efficient UI grounding. Guided by practical analysis of the task's characteristics and challenges, we propose FOCUSUI, an efficient UI grounding framework that selects patches most relevant to the instruction, while preserving positional continuity for precise grounding. FO-CUSUI addresses two key challenges: (1) Eliminating redundant tokens in visual encoding. We construct patchlevel supervision by fusing an instruction-conditioned and a rule-based UI-graph score that down-weights large homogeneous regions to select distinct and instruction-relevant visual tokens. (2) Preserving positional continuity during visual token selection. We find that general visual token pruning methods suffer from severe accuracy degradation on UI grounding tasks due to breaking positional information. We introduce a novel POSPAD strategy, which compresses each contiguous sequence of dropped visual tokens into a single special marker placed at the sequence's last index to preserve positional continuity. Comprehensive experiments on four grounding benchmarks demonstrate that FOCUSUI surpasses GUI-specific baselines. On the ScreenSpot-Pro benchmark, FOCUSUI-7B achieves performance improvement of 3.7% over GUI-Actor-7B. Also, even with only 30% visual token retention, the performance of FOCUSUI-7B only drops by 3.2%, while achieving up to 1.44× faster inference and 17% lower peak GPU memory. Decoding Decoding Decoding (a) Comparison of vanilla UI grounding VLMs, VLMs with visual token pruning, and our FOCUSUI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c41b9fe-bca5-4ddf-9418-760c8220d54dBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsQianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang et al.NeurIPS 2025 · 98 citations
- Token Merging: Your ViT But FasterDaniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang et al.ICLR 2023 · 62 citations
Related papers
- Visual Test-Time Scaling for GUI Agent GroundingTiange Luo, Lajanugen Logeswaran, Justin Johnson, Honglak LeeICCV 2025 · 3 citations
- ShowUI: One Vision-Language-Action Model for GUI Visual AgentKevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang et al.CVPR 2025
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu et al.CVPR 2026 · 3 citations
- SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language ModelsTong Xu, Hailong Shi, Xingyu GaoCVPR 2026
- Nüwa: Mending the Spatial Integrity Torn by VLM Token PruningYihong Huang, Fei Ma, Yihua Shao, Jingcai Guo et al.ICLR 2026 · 15 citations
