Visual Test-Time Scaling for GUI Agent Grounding
Tiange Luo, Lajanugen Logeswaran, Justin Johnson, Honglak Lee
Abstract
We introduce RegionFocus, a visual test-time scaling approach for Vision Language Model Agents. Understanding webpages is challenging due to the visual complexity of GUI images and the large number of interface elements, making accurate action selection difficult. Our approach dynamically zooms in on relevant regions, reducing background clutter and improving grounding accuracy. To support this process, we propose an image-as-map mechanism that visualizes key landmarks at each step, providing a transparent action record and enables the agent to effectively choose among action candidates. Even with a simple region selection strategy, we observe significant performance gains of 28+% on ScreenSpot-Pro and 24+% on WebVoyager benchmarks on top of two state-of-the-art open vision language model agents, UI-TARS-72B and Qwen2.5-VL-72B, highlighting the effectiveness of visual test-time scaling in interactive settings. We achieve a new state-of-the-art grounding performance of 61.6% on the ScreenSpot-Pro benchmark by applying RegionFocus to a Qwen2.5-VL-72B model. Our code is publicly available at https://github.com/tiangeluo/RegionFocus.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95685a0a-353a-4684-81e4-4ddf8b94aac6Cited by top-tier papers5
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual ReasoningHang Wu, Hongkai Chen, Yujun Cai, Chang Liu et al.EMNLP 2025 · 23 citations
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied AgentsWanxin Tian, Shijie Zhang, Kevin Zhang, Xiaowei Chi et al.NeurIPS 2025 · 20 citations
- MVP: Multiple View Prediction Improves GUI GroundingYunzhu Zhang, Zeyu Pan, Zhengwen Zeng, Shuheng Shen et al.CVPR 2026 · 10 citations
- Learning GUI Grounding with Spatial Reasoning from Visual FeedbackYu Zhao, Wei-Ning Chen, Huseyin Inan, Samuel Kessler et al.ICML 2026 · 9 citations
- Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI GroundingWenkai Wang, Xiyun Li, Hongcan Guo, Wenhao Yu et al.ACL 2026 · 1 citation
Builds on17
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 539 citations
Related papers
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu et al.CVPR 2026 · 3 citations
- FocusUI: Efficient UI Grounding via Position-Preserving Visual Token SelectionMingyu Ouyang, Kevin Qinghong Lin, Mike Zheng Shou, Hwee Tou NgCVPR 2026 · 8 citations
- GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsQianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang et al.NeurIPS 2025 · 98 citations
- Test-Time Reinforcement Learning for GUI Grounding via Region ConsistencyYong Du, Yuchen Yan, Fei Tang, Zhengxi Lu et al.AAAI 2026 · 10 citations
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual GroundingBin Lei, Nuo Xu, Ali Payani, Mingyi Hong et al.ICML 2026 · 8 citations
