GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
Bin Lei, Nuo Xu, Ali Payani, Mingyi Hong, Chunhua Liao, Yu Cao, Caiwen Ding
Abstract
Multimodal large language models (MLLMs) have markedly expanded the competence of graphical user-interface (GUI) systems, propelling them beyond controlled simulations into complex, real-world environments across diverse platforms. However, practical usefulness is still bounded by the reliability of visual grounding, i.e., mapping textual references to exact on-screen elements. This limitation prevents the system from accurately performing pointer-level actions such as clicking or dragging. To address it, we introduce GUI-SPOTLIGHT-A model trained for image-grounded reasoning that dynamically invokes multiple specialized tools to iteratively narrow its focus to the relevant region of the screen, thereby substantially improving visual grounding accuracy. On the ScreenSpot-Pro benchmark, GUI-Spotlight trained with only 18.5K training samples achieves 52.8% accuracy, surpassing V2P-7B (50.6% with 9.6M training samples) and GTA-1-7B (50.1% with 1.56M training samples). Code is avaliable at https://github.com/bin123apple/GUI_Spotlight .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25c3c4d6-970e-4824-8f45-6b85a31964dcCited by top-tier papers1
Ask how each one uses itBuilds on10
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang et al.ICLR 2026 · 109 citations
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain et al.NeurIPS 2025 · 90 citations
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo et al.ACM MM 2025 · 24 citations
Related papers
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu et al.CVPR 2026 · 3 citations
- Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI AgentsBoyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie et al.ICLR 2025
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement LearningWeitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding et al.ICLR 2026 · 6 citations
- GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI AgentsChen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu et al.AAAI 2026 · 5 citations
