WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics
Jason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng, Jeffrey Nichols, Jeffrey P. Bigham
Abstract
Modeling user interfaces (UIs) from visual information allows systems to make inferences about the functionality and semantics needed to support use cases in accessibility, app automation, and testing. Current datasets for training machine learning models are limited in size due to the costly and time-consuming process of manually collecting and annotating UIs. We crawled the web to construct WebUI, a large dataset of 400,000 rendered web pages associated with automatically extracted metadata. We analyze the composition of WebUI and show that while automatically extracted data is noisy, most examples meet basic criteria for visual UI modeling. We applied several strategies for incorporating semantics found in web pages to increase the performance of visual UI understanding models in the mobile domain, where less labeled data is available: (i) element detection, (ii) screen classification and (iii) screen similarity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49b395cc-4be6-471e-a0ed-e3939704e3d5Cited by top-tier papers26
- WebLINX: Real-World Website Navigation with Multi-Turn DialogueXing Han Lù, Zdenek Kasner, Siva ReddyICML 2024 · 146 citations
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li et al.ICLR 2026 · 54 citations
- UIClip: A Data-driven Model for Assessing User Interface DesignJason Wu, Yi-Hao Peng, Xin Yue Amanda Li, Amanda Swearngin et al.UIST 2024 · 29 citations
- Automating the Enterprise with Foundation ModelsMichael Wornow, Avanika Narayan, Krista Opsahl-Ong, Quinn McIntyre et al.VLDB 2024 · 26 citations
- CodeA11y: Making AI Coding Assistants Useful for Accessible Web DevelopmentPeya Mowar, Yi-Hao Peng, Jason Wu, Aaron Steinfeld et al.CHI 2025 · 22 citations
Builds on13
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White et al.CHI 2021 · 145 citations
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- VINS: Visual Search for Mobile User Interface DesignSara Bunian, Kai Li, Chaima Jemmali, Casper Harteveld et al.CHI 2021 · 100 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
Related papers
- Never-ending Learning of User InterfacesJason Wu, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin et al.UIST 2023 · 17 citations
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 62 citations
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu et al.ICLR 2025
- MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI ModelingSidong Feng, Suyu Ma, Han Wang, David Kong et al.CHI 2024 · 13 citations
- AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsHongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen et al.ACL 2025
