UICrit: Enhancing Automated Design Evaluation with a UI Critique Dataset
Peitong Duan, Chin-Yi Cheng, Gang Li, Bjoern Hartmann, Yang Li
Abstract
Automated UI evaluation can be beneficial for the design process; for example, to compare different UI designs, or conduct automated heuristic evaluation. LLM-based UI evaluation, in particular, holds the promise of generalizability to a wide variety of UI types and evaluation tasks. However, current LLM-based techniques do not yet match the performance of human evaluators. We hypothesize that automatic evaluation can be improved by collecting a targeted UI feedback dataset and then using this dataset to enhance the performance of general-purpose LLMs. We present a targeted dataset of 3,059 design critiques and quality ratings for 983 mobile UIs, collected from seven experienced designers. We carried out an in-depth analysis to characterize the dataset's features. We then applied this dataset to achieve a 55% performance gain in LLM-generated UI feedback via various few-shot and visual prompting techniques. We also discuss future applications of this dataset, including training a reward model for generative UI techniques, and fine-tuning a tool-agnostic multi-modal LLM that automates UI evaluation.
• Human-centered computing → Systems and tools for interaction design; Empirical studies in HCI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext facafbe8-710e-4dc0-a19a-af8505b11b43Cited by top-tier papers8
- ScreenAudit: Detecting Screen Reader Accessibility Errors in Mobile Apps Using Large Language ModelsMingyuan Zhong, Ruolin Chen, Xia Chen, James Fogarty et al.CHI 2025 · 15 citations
- SQUIRE: Interactive UI Authoring via Slot QUery Intermediate REpresentationsAlan Leung, Ruijia Cheng, Jason Wu, Jeffrey Nichols et al.UIST 2025 · 4 citations
- CritiqueCrew: Orchestrating Multi-Perspective Conversational Design CritiqueXiaojiao Chen, Jiahuan Zhou, Yunfeng Shu, Ruihan Wang et al.CHI 2026 · 2 citations
- VizCrit: Exploring Strategies for Displaying Computational Feedback in a Visual Design ToolMingyi Li, Mengyi Chen, Sarah Luo, Yining Cao et al.CHI 2026 · 1 citation
- Improving User Interface Generation Models from Designer FeedbackJason Wu, Amanda Swearngin, Arun Krishnavajjala, Alan Leung et al.CHI 2026 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel et al.CHI 2023 · 100 citations
- Generating Automatic Feedback on UI Mockups with Large Language ModelsPeitong Duan, Jeremy Warner, Yang Li, Bjoern HartmannCHI 2024 · 81 citations
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang et al.ACL 2020 · 75 citations
Related papers
- AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsHongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen et al.ACL 2025
- MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI ModelingSidong Feng, Suyu Ma, Han Wang, David Kong et al.CHI 2024 · 13 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation SlidesZhuohao Jerry Zhang, Ruiqi Chen, Mingyuan Zhong, Jacob O. WobbrockUIST 2025
- HypoEval: Hypothesis-Guided Evaluation for Natural Language GenerationMingxuan Li, Hanchen Li, Chenhao TanACL 2026 · 1 citation
