Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, Yang Li
Abstract
Mobile User Interface Summarization generates succinct language descriptions of mobile screens for conveying important contents and functionalities of the screen, which can be useful for many language-based application scenarios. We present Screen2Words, a novel screen summarization approach that automatically encapsulates essential information of a UI screen into a coherent language phrase. Summarizing mobile screens requires a holistic understanding of the multi-modal data of mobile UIs, including text, image, structures as well as UI semantics, motivating our multi-modal learning approach. We collected and analyzed a large-scale screen summarization dataset annotated by human workers. Our dataset contains more than 112k language summarization across ∼ 22k unique UI screens. We then experimented with a set of deep models with different configurations. Our evaluation of these models with both automatic accuracy metrics and human rating shows that our approach can generate high-quality summaries for mobile screens. We demonstrate potential use cases of Screen2Words and open-source our dataset and model to lay the foundations for further bridging language and user interfaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers67
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceTao Lei, Junwen Bai, Siddhartha Brahma, Joshua Ainslie et al.NeurIPS 2023 · 103 citations
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras et al.NeurIPS 2025 · 97 citations
Builds on8
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White et al.CHI 2021 · 145 citations
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li et al.AAAI 2020 · 113 citations
- VINS: Visual Search for Mobile User Interface DesignSara Bunian, Kai Li, Chaima Jemmali, Casper Harteveld et al.CHI 2021 · 100 citations
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang et al.ACL 2020 · 75 citations
Related papers
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng et al.EMNLP 2020 · 46 citations
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng et al.CHI 2023 · 49 citations
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu et al.ICLR 2025
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
- Screen2Vec: Semantic Embedding of GUI Screens and GUI ComponentsToby Jia-Jun Li, Lindsay Popowski, Tom M. Mitchell, Brad A. MyersCHI 2021 · 72 citations
