WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics
Jason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng, Jeffrey Nichols, Jeffrey P. Bigham
摘要
Modeling user interfaces (UIs) from visual information allows systems to make inferences about the functionality and semantics needed to support use cases in accessibility, app automation, and testing. Current datasets for training machine learning models are limited in size due to the costly and time-consuming process of manually collecting and annotating UIs. We crawled the web to construct WebUI, a large dataset of 400,000 rendered web pages associated with automatically extracted metadata. We analyze the composition of WebUI and show that while automatically extracted data is noisy, most examples meet basic criteria for visual UI modeling. We applied several strategies for incorporating semantics found in web pages to increase the performance of visual UI understanding models in the mobile domain, where less labeled data is available: (i) element detection, (ii) screen classification and (iii) screen similarity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- WebLINX: Real-World Website Navigation with Multi-Turn DialogueXing Han Lù, Zdenek Kasner, Siva ReddyICML 2024 · 被引用 146 次
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li 等ICLR 2026 · 被引用 54 次
- UIClip: A Data-driven Model for Assessing User Interface DesignJason Wu, Yi-Hao Peng, Xin Yue Amanda Li, Amanda Swearngin 等UIST 2024 · 被引用 29 次
- Automating the Enterprise with Foundation ModelsMichael Wornow, Avanika Narayan, Krista Opsahl-Ong, Quinn McIntyre 等VLDB 2024 · 被引用 26 次
- CodeA11y: Making AI Coding Assistants Useful for Accessible Web DevelopmentPeya Mowar, Yi-Hao Peng, Jason Wu, Aaron Steinfeld 等CHI 2025 · 被引用 22 次
它引用的顶会 Paper13
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White 等CHI 2021 · 被引用 145 次
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen 等FSE 2020 · 被引用 144 次
- VINS: Visual Search for Mobile User Interface DesignSara Bunian, Kai Li, Chaima Jemmali, Casper Harteveld 等CHI 2021 · 被引用 100 次
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen 等UIST 2021 · 被引用 97 次
相关 Paper
- Never-ending Learning of User InterfacesJason Wu, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin 等UIST 2023 · 被引用 17 次
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 被引用 62 次
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu 等ICLR 2025
- MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI ModelingSidong Feng, Suyu Ma, Han Wang, David Kong 等CHI 2024 · 被引用 13 次
- AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsHongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen 等ACL 2025
