Generating Automatic Feedback on UI Mockups with Large Language Models
Peitong Duan, Jeremy Warner, Yang Li, Bjoern Hartmann
Abstract
Feedback on user interface (UI) mockups is crucial in design. However, human feedback is not always readily available. We explore the potential of using large language models for automatic feedback. Specifically, we focus on applying GPT-4 to automate heuristic evaluation, which currently entails a human expert assessing a UI’s compliance with a set of design guidelines. We implemented a Figma plugin that takes in a UI design and a set of written heuristics, and renders automatically-generated feedback as constructive suggestions. We assessed performance on 51 UIs using three sets of guidelines, compared GPT-4-generated design suggestions with those from human experts, and conducted a study with 12 expert designers to understand fit with existing practice. We found that GPT-4-based feedback is useful for catching subtle errors, improving text, and considering UI semantics, but feedback also decreased in utility over iterations. Participants described several uses for this plugin despite its imperfect suggestions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- UIClip: A Data-driven Model for Assessing User Interface DesignJason Wu, Yi-Hao Peng, Xin Yue Amanda Li, Amanda Swearngin et al.UIST 2024 · 29 citations
- LogoMotion: Visually-Grounded Code Synthesis for Creating and Editing AnimationVivian Liu, Rubaiat Habib Kazi, Li-Yi Wei, Matthew Fisher et al.CHI 2025 · 23 citations
- UICrit: Enhancing Automated Design Evaluation with a UI Critique DatasetPeitong Duan, Chin-Yi Cheng, Gang Li, Bjoern Hartmann et al.UIST 2024 · 22 citations
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Closing the Loop between User Stories and GUI Prototypes: An LLM-Based Assistant for Cross-Functional Integration in Software DevelopmentFelix Kretzer, Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto et al.CHI 2025 · 18 citations
- DesignRepair: Dual-Stream Design Guideline-Aware Frontend Repair with Large Language ModelsMingyue Yuan, Jieshan Chen, Zhenchang Xing, Aaron Quigley et al.ICSE 2025 · 2 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- "Create a Fear of Missing Out" - ChatGPT Implements Unsolicited Deceptive Designs in Generated Websites Without WarningVeronika Krauß, Mark McGill, Thomas Kosch, Yolanda Maira Thiel et al.CHI 2025 · 17 citations
- SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile ApplicationsWei Xiang, Hanfei Zhu, Suqi Lou, Xinli Chen et al.CHI 2024 · 49 citations
