MVP: Multiple View Prediction Improves GUI Grounding
Yunzhu Zhang, Zeyu Pan, Zhengwen Zeng, Shuheng Shen, Changhua Meng, Linchao Zhu
Abstract
GUI grounding, which translates natural language instructions into precise pixel coordinates, is essential for developing practical GUI agents. However, we observe that existing grounding models exhibit significant coordinate prediction instability-minor visual perturbations (e.g., cropping a few pixels) can drastically alter predictions, flipping results between correct and incorrect. This instability severely undermines model performance, especially for samples with high-resolution and small UI elements. To address this issue, we propose Multi-View Prediction (MVP), a training-free framework that enhances grounding performance through multi-view inference. Our key insight is that while single-view predictions may be unstable, aggregating predictions from multiple carefully cropped views can effectively distinguish correct coordinates from outliers. MVP comprises two components: (1) Attention-Guided View Proposal, which derives diverse views guided by instruction-to-image attention scores, and (2) Multi-Coordinates Clustering, which ensembles predictions by selecting the centroid of the densest spatial cluster. Extensive experiments demonstrate MVP's effectiveness across various models and benchmarks. Notably, on ScreenSpot-Pro, MVP boosts UI-TARS-1.5-7B to 56.1%, GTA1-7B to 61.7%, Qwen3VL-8B-Instruct to 65.3%, and Qwen3VL-32B-Instruct to 74.0%. The code is available at https: //github.com/ZJUSCL/MVP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28332470-5994-4538-b68b-b9f0b6d8aceaBuilds on17
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang et al.ICLR 2026 · 109 citations
- GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsQianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang et al.NeurIPS 2025 · 98 citations
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo et al.ACM MM 2025 · 24 citations
Related papers
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu et al.CVPR 2026 · 3 citations
- BAMI: Training-Free Bias Mitigation in GUI GroundingBorui Zhang, Bo Zhang, Bo Wang, Wenzhao Zheng et al.CVPR 2026 · 2 citations
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as ReasoningLiangyu Chen, Hanzhang Zhou, chenglin Cai, Jianan Zhang et al.ICLR 2026 · 19 citations
- Test-Time Reinforcement Learning for GUI Grounding via Region ConsistencyYong Du, Yuchen Yan, Fei Tang, Zhengxi Lu et al.AAAI 2026 · 10 citations
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual ReasoningHang Wu, Hongkai Chen, Yujun Cai, Chang Liu et al.EMNLP 2025 · 23 citations
