Psychologically-inspired, unsupervised inference of perceptual groups of GUI widgets from GUI images
Mulong Xie, Zhenchang Xing, Sidong Feng, Xiwei Xu, Liming Zhu, Chunyang Chen
Abstract
Graphical User Interface (GUI) is not merely a collection of individual and unrelated widgets, but rather partitions discrete widgets into groups by various visual cues, thus forming higher-order perceptual units such as tab, menu, card or list. The ability to automatically segment a GUI into perceptual groups of widgets constitutes a fundamental component of visual intelligence to automate GUI design, implementation and automation tasks. Although humans can partition a GUI into meaningful perceptual groups of widgets in a highly reliable way, perceptual grouping is still an open challenge for computational approaches. Existing methods rely on ad-hoc heuristics or supervised machine learning that is dependent on specific GUI implementations and runtime information. Research in psychology and biological vision has formulated a set of principles (i.e., Gestalt theory of perception) that describe how humans group elements in visual scenes based on visual cues like connectivity, similarity, proximity and continuity. These principles are domain-independent and have been widely adopted by practitioners to structure content on GUIs to improve aesthetic pleasantness and usability. Inspired by these principles, we present a novel unsupervised image-based method for inferring perceptual groups of GUI widgets. Our method requires only GUI pixel images, is independent of GUI implementation, and does not require any training data. The evaluation on a dataset of 1,091 GUIs collected from 772 mobile apps and 20 UI design mockups shows that our method significantly outperforms the state-of-the-art ad-hoc heuristics-based baseline. Our perceptual grouping method creates opportunities for improving UI-related software engineering tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 143 citations
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningYunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma et al.UIST 2024 · 24 citations
- Efficiency Matters: Speeding Up Automated Testing with GUI Rendering InferenceSidong Feng, Mulong Xie, Chunyang ChenICSE 2023 · 23 citations
- Feedback-Driven Automated Whole Bug Report Reproduction for Android AppsDingbang Wang, Yu Zhao, Sidong Feng, Zhaoxu Zhang et al.ISSTA 2024 · 16 citations
- MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI ModelingSidong Feng, Suyu Ma, Han Wang, David Kong et al.CHI 2024 · 13 citations
Builds on8
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White et al.CHI 2021 · 145 citations
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- Unblind your apps: predicting natural-language labels for mobile GUI components by deep learningJieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu et al.ICSE 2020 · 101 citations
- From Lost to Found: Discover Missing UI Design Semantics through Recovering Missing TagsChunyang Chen, Sidong Feng, Zhengyang Liu, Zhenchang Xing et al.CSCW 2020 · 41 citations
- GIFdroid: Automated Replay of Visual Bug Reports for Android AppsSidong Feng, Chunyang ChenICSE 2022 · 38 citations
Related papers
- Latent Noise Segmentation: How Neural Noise Leads to the Emergence of Segmentation and GroupingBen Lonnqvist, Zhengqing Wu, Michael H. HerzogICML 2024 · 5 citations
- Slide Gestalt: Automatic Structure Extraction in Slide Decks for Non-Visual AccessYi-Hao Peng, Peggy Chi, Anjuli Kannan, Meredith Ringel Morris et al.CHI 2023 · 18 citations
- Graph4GUI: Graph Neural Networks for Representing Graphical User InterfacesYue Jiang, Changkong Zhou, Vikas Garg, Antti OulasvirtaCHI 2024 · 17 citations
- GUIGAN: Learning to Generate GUI Designs Using Generative Adversarial NetworksTianming Zhao, Chunyang Chen, Yuanning Liu, Xiaodong ZhuICSE 2021 · 58 citations
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 62 citations
