CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
Daeheon Jeong, Seoyeon Byun, Kihoon Son, Dae Hyun Kim, Juho Kim
摘要
User interface (UI) design is an iterative process in which designers progressively refine their work with design software such as Figma or Sketch. Recent advances in vision–language models (VLMs) with tool invocation suggest these models can operate design software to edit a UI design through iteration. Understanding and enhancing this capacity is important, as it highlights VLMs’ potential to collaborate with designers within conventional software. However, as no existing benchmark evaluates tool-based design performance, the capacity remains unknown. To address this, we introduce CANVAS, a benchmark for VLMs on tool-based user interface design. Our benchmark contains 598 tool-based design tasks paired with ground-truth references sampled from 3.3K mobile UI designs across 30 function-based categories (e.g., onboarding, messaging). In each task, a VLM updates the design step-by-step through context-based tool invocations (e.g., create a rectangle as a button background), linked to design software. Specifically, CANVAS incorporates two task types: (i) design replication evaluates the ability to reproduce a whole UI screen; (ii) design modification evaluates the ability to modify a specific part of an existing screen. Results suggest that leading models exhibit more strategic tool invocations, improving design quality. Furthermore, we identify common error patterns models exhibit, guiding future work in enhancing tool-based design capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- User Experience Design Professionals' Perceptions of Generative Artificial IntelligenceJie Li, Hancheng Cao, Laura Lin, Youyang Hou 等CHI 2024 · 被引用 149 次
- CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AIDaEun Choi, Sumin Hong, Jeongeon Park, John Joon Young Chung 等CHI 2024 · 被引用 116 次
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel 等CHI 2023 · 被引用 100 次
- Stylette: Styling the Web with Natural LanguageTae Soo Kim, DaEun Choi, Yoonseo Choi, Juho KimCHI 2022 · 被引用 71 次
相关 Paper
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design UnderstandingJaehyun Jeon, Min Soo Kim, Janghan Yoon, Sumin Shim 等ACL 2026 · 被引用 1 次
- Canvil: Designerly Adaptation for LLM-Powered User ExperiencesK. J. Kevin Feng, Q. Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan 等CHI 2025 · 被引用 14 次
- Closing the Loop between User Stories and GUI Prototypes: An LLM-Based Assistant for Cross-Functional Integration in Software DevelopmentFelix Kretzer, Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto 等CHI 2025 · 被引用 18 次
- Co-Constructed or Constrained? How AI Collaboration Tools Reshape UI Design Practice in a Time-Boxed Design ChallengeCharlotte Kobiella, Lukas Schneider, Albrecht Schmidt, Nada TerzimehicCHI 2026 · 被引用 1 次
- Figma2Code: Automating Multimodal Design to Code in the WildYi Gui, Jiawan Zhang, Yina Wang, Tianran Ma 等ICLR 2026 · 被引用 3 次
