ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces
Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby B. Lee, Jindong Chen
Abstract
As mobile devices are becoming ubiquitous, regularly interacting with a variety of user interfaces (UIs) is a common aspect of daily life for many people. To improve the accessibility of these devices and to enable their usage in a variety of settings, building models that can assist users and accomplish tasks through the UI is vitally important. However, there are several challenges to achieve this. First, UI components of similar appearance can have different functionalities, making understanding their function more important than just analyzing their appearance. Second, domain-specific features like Document Object Model (DOM) in web pages and View Hierarchy (VH) in mobile applications provide important signals about the semantics of UI elements, but these features are not in a natural language format. Third, owing to a large diversity in UIs and absence of standard DOM or VH representations, building a UI understanding model with high coverage requires large amounts of training data.
Inspired by the success of pre-training based approaches in NLP for tackling a variety of problems in a data-efficient way, we introduce a new pre-trained UI representation model called ActionBert. Our methodology is designed to leverage visual, linguistic and domain-specific features in user interaction traces to pre-train generic feature representations of UIs and their components. Our key intuition is that user actions, e.g., a sequence of clicks on different UI components, reveals important information about their functionality. We evaluate the proposed model on a wide variety of downstream tasks, ranging from icon classification to UI component retrieval based on its natural language description. Experiments show that the proposed ActionBert model outperforms multi-modal baselines across all downstream tasks by up to 15.5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b2adfa1-3277-4554-95fc-36ee87e93647Cited by top-tier papers19
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo et al.ICLR 2024 · 160 citations
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng et al.CHI 2023 · 49 citations
- Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLMZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.CHI 2024 · 29 citations
- Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency AnalysisEldon Schoop, Xin Zhou, Gang Li, Zhourong Chen et al.CHI 2022 · 29 citations
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningYunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma et al.UIST 2024 · 24 citations
Builds on2
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Predicting and Diagnosing User Engagement with Mobile UI Animation via a Data-Driven ApproachZiming Wu, Yulun Jiang, Yiding Liu, Xiaojuan MaCHI 2020 · 50 citations
Related papers
- U-BERT: Pre-training User Representations for Improved RecommendationZhaopeng Qiu, Xian Wu, Jingyue Gao, Wei FanAAAI 2021 · 171 citations
- UIPro: Unleashing Superior Interaction Capability for GUI AgentsHongxin Li, Jingran Su, Jingfan Chen, Zheng Ju et al.ICCV 2025
- Spotlight: Mobile UI Understanding using Vision-Language Models with a FocusGang Li, Yang LiICLR 2023
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- Ferret-UI 2: Mastering Universal User Interface Understanding Across PlatformsZhangheng Li, Keen You, Haotian Zhang, Di Feng et al.ICLR 2025
