Figma2Code: Automating Multimodal Design to Code in the Wild
Yi Gui, Jiawan Zhang, Yina Wang, Tianran Ma, Yao Wan, Shilin He, Dongping Chen, Zhou Zhao, Wenbin Jiang, Xuanhua Shi, Hai Jin, Philip S. Yu
摘要
Front-end development constitutes a substantial portion of software engineering, yet converting design mockups into production-ready User Interface (UI) code remains tedious and time-costly.
While recent work has explored automating this process with Multimodal Large Language Models (MLLMs), existing approaches typically rely solely on design images. As a result, they must infer complex UI details from images alone, often leading to degraded results.
In real-world development workflows, however, design mockups are usually delivered as Figma files—a widely used tool for front-end design—that embed rich multimodal information (e.g., metadata and assets) essential for generating high-quality UI.
To bridge this gap, we introduce Figma2Code, a new task that generalizes design-to-code into a multimodal setting and aims to automate design-to-code in the wild.
Specifically, we collect paired design images and their corresponding metadata files from the Figma community. We then apply a series of processing operations, including rule-based filtering, human and MLLM-based annotation and screening, and metadata refinement. This process yields 3,055 samples, from which designers curate a balanced dataset of 213 high-quality cases.
Using this dataset, we benchmark ten state-of-the-art open-source and proprietary MLLMs. Our results show that while proprietary models achieve superior visual fidelity, they remain limited in layout responsiveness and code maintainability.
Further experiments across modalities and ablation studies corroborate this limitation, partly due to models’ tendency to directly map primitive visual attributes from Figma metadata.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 被引用 62 次
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng 等CHI 2023 · 被引用 49 次
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi 等WWW 2025 · 被引用 38 次
相关 Paper
- MLLM-Based UI2Code Automation Guided by UI Layout InformationFan Wu, Cuiyun Gao, Shuqing Li, Xin-Cheng Wen 等ISSTA 2025 · 被引用 5 次
- Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive PrototypingJingyu Xiao, Yuxuan Wan, Yintong Huo, Zixin Wang 等ASE 2025 · 被引用 1 次
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsHouston H. Zhang, Tao Zhang, Baoze Lin, Yuanqi Xue 等CVPR 2026 · 被引用 9 次
- Closing the Loop between User Stories and GUI Prototypes: An LLM-Based Assistant for Cross-Functional Integration in Software DevelopmentFelix Kretzer, Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto 等CHI 2025 · 被引用 18 次
- Divide-and-Conquer: Generating UI Code from ScreenshotsYuxuan Wan, Chaozheng Wang, Yi Dong, Wenxuan Wang 等FSE 2025 · 被引用 10 次
