DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
Hang Wu, Hongkai Chen, Yujun Cai, Chang Liu, Qingwen Ye, Ming-Hsuan Yang, Yiwei Wang
Abstract
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model's initial predictions and incrementally zooms into subregions to refine the grounding result. This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations. We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy OptimizationYuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li et al.AAAI 2026 · 15 citations
- Test-Time Reinforcement Learning for GUI Grounding via Region ConsistencyYong Du, Yuchen Yan, Fei Tang, Zhengxi Lu et al.AAAI 2026 · 10 citations
- GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI AgentsChen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu et al.AAAI 2026 · 5 citations
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying TogglesZongru Wu, Rui Mao, Zhiyuan Tian, Pengzhou Cheng et al.CVPR 2026 · 3 citations
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu et al.CVPR 2026 · 3 citations
Builds on8
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou et al.NeurIPS 2025 · 73 citations
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMsPenghao Wu, Saining XieCVPR 2024 · 32 citations
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo et al.ACM MM 2025 · 24 citations
- Visual Test-Time Scaling for GUI Agent GroundingTiange Luo, Lajanugen Logeswaran, Justin Johnson, Honglak LeeICCV 2025 · 3 citations
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su et al.CVPR 2024
Related papers
- Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal FusionLonghui Ma, Di Zhao, Siwei Wang, Zhao Lv et al.ICML 2026 · 3 citations
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-TuningHai-Ming Xu, Qi Chen, Lei Wang, Lingqiao LiuAAAI 2025 · 13 citations
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual GroundingBin Lei, Nuo Xu, Ali Payani, Mingyi Hong et al.ICML 2026 · 8 citations
- MVP: Multiple View Prediction Improves GUI GroundingYunzhu Zhang, Zeyu Pan, Zhengwen Zeng, Shuheng Shen et al.CVPR 2026 · 10 citations
