MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
Sidong Feng, Suyu Ma, Han Wang, David Kong, Chunyang Chen
摘要
The importance of computational modeling of mobile user interfaces (UIs) is undeniable. However, these require a high-quality UI dataset. Existing datasets are often outdated, collected years ago, and are frequently noisy with mismatches in their visual representation. This presents challenges in modeling UI understanding in the wild. This paper introduces a novel approach to automatically mine UI data from Android apps, leveraging Large Language Models (LLMs) to mimic human-like exploration. To ensure dataset quality, we employ the best practices in UI noise filtering and incorporate human annotation as a final validation step. Our results demonstrate the effectiveness of LLMs-enhanced app exploration in mining more meaningful UIs, resulting in a large dataset MUD of 18k human-annotated UIs from 3.3k apps. We highlight the usefulness of MUD in two common UI modeling tasks: element detection and UI retrieval, showcasing its potential to establish a foundation for future research into high-quality, modern UIs.
• Human-centered computing → Human computer interaction (HCI).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 被引用 33 次
- AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based AgentsJunting Lu, Zhiyang Zhang, Fangkai Yang, Jue Zhang 等ACL 2025 · 被引用 9 次
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun 等CHI 2026 · 被引用 2 次
- Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature TestingSidong Feng, Changhao Du, Huaxiao Liu, Qingnan Wang 等ICSE 2026 · 被引用 2 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White 等CHI 2021 · 被引用 145 次
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen 等FSE 2020 · 被引用 144 次
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 被引用 143 次
相关 Paper
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng 等CHI 2023 · 被引用 49 次
- Learning to Denoise Raw Mobile UI Layouts for Improving Datasets at ScaleGang Li, Gilles Baechler, Manuel Tragut, Yang LiCHI 2022 · 被引用 31 次
- UICrit: Enhancing Automated Design Evaluation with a UI Critique DatasetPeitong Duan, Chin-Yi Cheng, Gang Li, Bjoern Hartmann 等UIST 2024 · 被引用 22 次
- AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsHongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen 等ACL 2025
- LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task AutomationLi Zhang, Shihe Wang, Xianqing Jia, Zhihan Zheng 等UIST 2024 · 被引用 8 次
