Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning
Hai-Ming Xu, Qi Chen, Lei Wang, Lingqiao Liu
摘要
Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding-accurately identifying critical GUI components such as text or icons based on a GUI image and a corresponding text query. Traditionally, this task has relied on fine-tuning MLLMs with specialized training data to predict component locations directly. However, in this paper, we propose a novel Tuning-free Attention-driven Grounding (TAG) method that leverages the inherent attention patterns in pretrained MLLMs to accomplish this task without the need for additional fine-tuning. Our method involves identifying and aggregating attention maps from specific tokens within a carefully constructed query prompt. Applied to MiniCPM-Llama3-V 2.5, a state-of-the-art MLLM, our tuning-free approach achieves performance comparable to tuning-based methods, with notable success in text localization. Additionally, we demonstrate that our attention map-based grounding technique significantly outperforms direct localization predictions from MiniCPM-Llama3-V 2.5, highlighting the potential of using attention maps from pretrained MLLMs and paving the way for future innovations in this domain. Code is available at https://github.com/HeimingX/TAG.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsQianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang 等NeurIPS 2025 · 被引用 98 次
- MVP: Multiple View Prediction Improves GUI GroundingYunzhu Zhang, Zeyu Pan, Zhengwen Zeng, Shuheng Shen 等CVPR 2026 · 被引用 10 次
- Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal FusionLonghui Ma, Di Zhao, Siwei Wang, Zhao Lv 等ICML 2026 · 被引用 3 次
- Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI GroundingWenkai Wang, Xiyun Li, Hongcan Guo, Wenhao Yu 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper10
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu 等ACL 2024 · 被引用 30 次
相关 Paper
- Enhancing Part-Level Point Grounding for Any Open-Source MLLMsJin-Cheng Jhang, Fu-En Wang, Xin Yang, Nan Qiao 等CVPR 2026
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual GroundingSeil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae HwangCVPR 2025
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual ReasoningHang Wu, Hongkai Chen, Yujun Cai, Chang Liu 等EMNLP 2025 · 被引用 23 次
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu 等CVPR 2026 · 被引用 3 次
