UICrit: Enhancing Automated Design Evaluation with a UI Critique Dataset
Peitong Duan, Chin-Yi Cheng, Gang Li, Bjoern Hartmann, Yang Li
摘要
Automated UI evaluation can be beneficial for the design process; for example, to compare different UI designs, or conduct automated heuristic evaluation. LLM-based UI evaluation, in particular, holds the promise of generalizability to a wide variety of UI types and evaluation tasks. However, current LLM-based techniques do not yet match the performance of human evaluators. We hypothesize that automatic evaluation can be improved by collecting a targeted UI feedback dataset and then using this dataset to enhance the performance of general-purpose LLMs. We present a targeted dataset of 3,059 design critiques and quality ratings for 983 mobile UIs, collected from seven experienced designers. We carried out an in-depth analysis to characterize the dataset's features. We then applied this dataset to achieve a 55% performance gain in LLM-generated UI feedback via various few-shot and visual prompting techniques. We also discuss future applications of this dataset, including training a reward model for generative UI techniques, and fine-tuning a tool-agnostic multi-modal LLM that automates UI evaluation.
• Human-centered computing → Systems and tools for interaction design; Empirical studies in HCI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ScreenAudit: Detecting Screen Reader Accessibility Errors in Mobile Apps Using Large Language ModelsMingyuan Zhong, Ruolin Chen, Xia Chen, James Fogarty 等CHI 2025 · 被引用 15 次
- SQUIRE: Interactive UI Authoring via Slot QUery Intermediate REpresentationsAlan Leung, Ruijia Cheng, Jason Wu, Jeffrey Nichols 等UIST 2025 · 被引用 4 次
- CritiqueCrew: Orchestrating Multi-Perspective Conversational Design CritiqueXiaojiao Chen, Jiahuan Zhou, Yunfeng Shu, Ruihan Wang 等CHI 2026 · 被引用 2 次
- VizCrit: Exploring Strategies for Displaying Computational Feedback in a Visual Design ToolMingyi Li, Mengyi Chen, Sarah Luo, Yining Cao 等CHI 2026 · 被引用 1 次
- Improving User Interface Generation Models from Designer FeedbackJason Wu, Amanda Swearngin, Arun Krishnavajjala, Alan Leung 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel 等CHI 2023 · 被引用 100 次
- Generating Automatic Feedback on UI Mockups with Large Language ModelsPeitong Duan, Jeremy Warner, Yang Li, Bjoern HartmannCHI 2024 · 被引用 81 次
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang 等ACL 2020 · 被引用 75 次
相关 Paper
- AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsHongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen 等ACL 2025
- MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI ModelingSidong Feng, Suyu Ma, Han Wang, David Kong 等CHI 2024 · 被引用 13 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
- SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation SlidesZhuohao Jerry Zhang, Ruiqi Chen, Mingyuan Zhong, Jacob O. WobbrockUIST 2025
- HypoEval: Hypothesis-Guided Evaluation for Natural Language GenerationMingxuan Li, Hanchen Li, Chenhao TanACL 2026 · 被引用 1 次
