Screen2Vec: Semantic Embedding of GUI Screens and GUI Components
Toby Jia-Jun Li, Lindsay Popowski, Tom M. Mitchell, Brad A. Myers
Abstract
Representing the semantics of GUI screens and components is crucial to data-driven computational methods for modeling user-GUI interactions and mining GUI designs. Existing GUI semantic representations are limited to encoding either the textual content, the visual design and layout patterns, or the app contexts. Many representation techniques also require significant manual data annotation efforts. This paper presents Screen2Vec, a new self-supervised technique for generating representations in embedding vectors of GUI screens and components that encode all of the above GUI features without requiring manual annotation using the context of user interaction traces. Screen2Vec is inspired by the word embedding method Word2Vec, but uses a new two-layer pipeline informed by the structure of GUIs and interaction traces and incorporates screen- and app-specific metadata. Through several sample downstream tasks, we demonstrate Screen2Vec’s key useful properties: representing between-screen similarity through nearest neighbors, composability, and capability to represent user tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee8e0df7-4a3c-45a4-af5e-77a317f76dd6Cited by top-tier papers39
- Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Literature ReviewYang Shi, Tian Gao, Xiaohan Jiao, Nan CaoCSCW 2023 · 170 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- CanvasVAE: Learning to Generate Vector Graphic DocumentsKota YamaguchiICCV 2021 · 103 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
Builds on6
- Unblind your apps: predicting natural-language labels for mobile GUI components by deep learningJieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu et al.ICSE 2020 · 101 citations
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang et al.ACL 2020 · 75 citations
- GUIComp: A GUI Design Assistant with Real-Time, Multi-Faceted FeedbackChunggi Lee, Sanghoon Kim, Dongyun Han, Hongjun Yang et al.CHI 2020 · 61 citations
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng et al.EMNLP 2020 · 46 citations
Related papers
- Mouse2Vec: Learning Reusable Semantic Representations of Mouse BehaviourGuanhua Zhang, Zhiming Hu, Mihai Bâce, Andreas BullingCHI 2024 · 5 citations
- Graph4GUI: Graph Neural Networks for Representing Graphical User InterfacesYue Jiang, Changkong Zhou, Vikas Garg, Antti OulasvirtaCHI 2024 · 17 citations
- CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic DesignYuxi Xie, Danqing Huang, Jinpeng Wang, Chin-Yew LinACM MM 2021 · 11 citations
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- GUIGAN: Learning to Generate GUI Designs Using Generative Adversarial NetworksTianming Zhao, Chunyang Chen, Yuanning Liu, Xiaodong ZhuICSE 2021 · 58 citations
