UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI Agents
Jiwen Zhang, Ya-Qi Yu, Minghui Liao, WenTao Li, Jihao Wu, Zhongyu Wei
Abstract
Graphical User Interface (GUI) agents are expected to precisely operate on the screens of digital devices. Existing GUI agents merely depend on current visual observations and plaintext action history, ignoring the significance of history screens. To mitigate this issue, we propose UI-Hawk, a multi-modal GUI agent specially designed to process screen streams encountered during GUI navigation. UI-Hawk incorporates a history-aware visual encoder to handle the screen sequences. To acquire a better understanding of screen streams, we select four fundamental tasks-UI grounding, UI referring, screen question answering, and screen summarization. We further propose a curriculum learning strategy to subsequently guide the model from fundamental tasks to advanced screen-stream comprehension. Along with the efforts above, we have created a benchmark FunUI to quantitatively evaluate the fundamental screen understanding ability of MLLMs. Extensive experiments on FunUI and GUI navigation benchmarks consistently validate that screen stream understanding is essential for GUI tasks. Our code and data are now available at https://github.com/IMNearth/UIHawk .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7d7e123-6f10-453d-a4c6-4088f6e0735bCited by top-tier papers1
Ask how each one uses itBuilds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White et al.CHI 2021 · 145 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
Related papers
- UIPro: Unleashing Superior Interaction Capability for GUI AgentsHongxin Li, Jingran Su, Jingfan Chen, Zheng Ju et al.ICCV 2025
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- GUI-Rise: Structured Reasoning and History Summarization for GUI NavigationTao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu et al.NeurIPS 2025 · 4 citations
- Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens GroundingYue Fan, Lei Ding, Ching-Chen Kuo, Shan Jiang et al.EMNLP 2024 · 1 citation
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
