Figma2Code: Automating Multimodal Design to Code in the Wild
Yi Gui, Jiawan Zhang, Yina Wang, Tianran Ma, Yao Wan, Shilin He, Dongping Chen, Zhou Zhao, Wenbin Jiang, Xuanhua Shi, Hai Jin, Philip S. Yu
Abstract
Front-end development constitutes a substantial portion of software engineering, yet converting design mockups into production-ready User Interface (UI) code remains tedious and time-costly.
While recent work has explored automating this process with Multimodal Large Language Models (MLLMs), existing approaches typically rely solely on design images. As a result, they must infer complex UI details from images alone, often leading to degraded results.
In real-world development workflows, however, design mockups are usually delivered as Figma files—a widely used tool for front-end design—that embed rich multimodal information (e.g., metadata and assets) essential for generating high-quality UI.
To bridge this gap, we introduce Figma2Code, a new task that generalizes design-to-code into a multimodal setting and aims to automate design-to-code in the wild.
Specifically, we collect paired design images and their corresponding metadata files from the Figma community. We then apply a series of processing operations, including rule-based filtering, human and MLLM-based annotation and screening, and metadata refinement. This process yields 3,055 samples, from which designers curate a balanced dataset of 213 high-quality cases.
Using this dataset, we benchmark ten state-of-the-art open-source and proprietary MLLMs. Our results show that while proprietary models achieve superior visual fidelity, they remain limited in layout responsiveness and code maintainability.
Further experiments across modalities and ablation studies corroborate this limitation, partly due to models’ tendency to directly map primitive visual attributes from Figma metadata.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1be538a7-a020-47a2-95c1-c454260ef2b8Builds on12
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 62 citations
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng et al.CHI 2023 · 49 citations
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
Related papers
- MLLM-Based UI2Code Automation Guided by UI Layout InformationFan Wu, Cuiyun Gao, Shuqing Li, Xin-Cheng Wen et al.ISSTA 2025 · 5 citations
- Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive PrototypingJingyu Xiao, Yuxuan Wan, Yintong Huo, Zixin Wang et al.ASE 2025 · 1 citation
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsHouston H. Zhang, Tao Zhang, Baoze Lin, Yuanqi Xue et al.CVPR 2026 · 9 citations
- Closing the Loop between User Stories and GUI Prototypes: An LLM-Based Assistant for Cross-Functional Integration in Software DevelopmentFelix Kretzer, Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto et al.CHI 2025 · 18 citations
- Divide-and-Conquer: Generating UI Code from ScreenshotsYuxuan Wan, Chaozheng Wang, Yi Dong, Wenxuan Wang et al.FSE 2025 · 10 citations
