Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
Yue Fan, Lei Ding, Ching-Chen Kuo, Shan Jiang, Yang Zhao, Xinze Guan, Jie Yang, Yi Zhang, Xin Wang
Abstract
Graphical User Interfaces (GUIs) are central to our interaction with digital devices and growing efforts have been made to build models for various GUI understanding tasks. However, these efforts largely overlook an important GUI-referring task: screen reading based on user-indicated points, which we name the Screen Point-and-Read (ScreenPR) task. Currently, this task is predominantly handled by rigid accessible screen reading tools, in great need of new models driven by advancements in Multimodal Large Language Models (MLLMs). In this paper, we propose a Tree-of-Lens (ToL) agent, utilizing a novel ToL grounding mechanism, to address the ScreenPR task. Based on the input point coordinate and the corresponding GUI screenshot, our ToL agent constructs a Hierarchical Layout Tree. Based on the tree, our ToL agent not only comprehends the content of the indicated area but also articulates the layout and spatial relationships between elements. Such layout information is crucial for accurately interpreting information on the screen, distinguishing our ToL agent from other screen reading tools. We also thoroughly evaluate the ToL agent against other baselines on a newly proposed ScreenPR benchmark, which includes GUIs from mobile, web, and operating systems. Last but not least, we test the ToL agent on mobile GUI navigation tasks, demonstrating its utility in identifying incorrect actions along the path of agent execution trajectories. Code and data: https://screen-point-and-read.github.io . Lens 1 with a fine-grained field width Lens 2 with a coarse field width Local region. Global region.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous ExplorationYue Fan, Handong Zhao, Ruiyi Zhang, Yu Shen et al.EMNLP 2025 · 15 citations
- Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-TuningHai-Ming Xu, Qi Chen, Lei Wang, Lingqiao LiuAAAI 2025 · 13 citations
- VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-useZhehao Zhang, Ryan A. Rossi, Tong Yu, Franck Dernoncourt et al.AAAI 2026 · 11 citations
- UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI AgentsJiwen Zhang, Ya-Qi Yu, Minghui Liao, WenTao Li et al.EMNLP 2025 · 1 citation
- OSCAR: Operating System Control via State-Aware Reasoning and Re-PlanningXiaoqiang Wang, Bang LiuICLR 2025
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
Related papers
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo et al.ACM MM 2025 · 24 citations
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu et al.CVPR 2026 · 3 citations
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- CogAgent: A Visual Language Model for GUI AgentsWenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu et al.CVPR 2024
