MLLM-Based UI2Code Automation Guided by UI Layout Information
Fan Wu, Cuiyun Gao, Shuqing Li, Xin-Cheng Wen, Qing Liao
Abstract
Converting user interfaces into code (UI2Code) is a crucial step in website development, which is time-consuming and labor-intensive. The automation of UI2Code is essential to streamline this task, beneficial for improving the development efficiency. There exist deep learning-based methods for the task; however, they heavily rely on a large amount of labeled training data and struggle with generalizing to real-world, unseen web page designs. The advent of Multimodal Large Language Models (MLLMs) presents potential for alleviating the issue, but they are difficult to comprehend the complex layouts in UIs and generate the accurate code with layout preserved. To address these issues, we propose LayoutCoder, a novel MLLM-based framework generating UI code from real-world webpage images, which includes three key modules: (1) Element Relation Construction, which aims at capturing UI layout by identifying and grouping components with similar structures; (2) UI Layout Parsing, which aims at generating UI layout trees for guiding the subsequent code generation process; and (3) Layout-Guided Code Fusion, which aims at producing the accurate code with layout preserved. For evaluation, we build a new benchmark dataset which involves 350 real-world websites named Snap2Code, divided into seen and unseen parts for mitigating the data leakage issue, besides the popular dataset Design2Code. Extensive evaluation shows the superior performance of LayoutCoder over the state-of-the-art approaches. Compared with the best-performing baseline, LayoutCoder improves 10.14% in the BLEU score and 3.95% in the CLIP score on average across all datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7900484-260a-4943-97cc-a202e66dbac2Cited by top-tier papers5
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsHouston H. Zhang, Tao Zhang, Baoze Lin, Yuanqi Xue et al.CVPR 2026 · 9 citations
- MulFCoder: Framework-conditioned Multi-agent for MLLM-based Multi-framework Front-end Code GenerationJie Wu, Haoran Ma, Shisong Tang, Yulin Xu et al.ICML 2026
- EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code GenerationJingyu Xiao, Zhongyi Zhang, Yuxuan Wan, Yintong Huo et al.FSE 2026
- Deterministic Component Mining for Multi-Framework UI2Code GenerationZixiong Yang, Linxiao Li, Jiaye Lin, Binrui Wu et al.ICML 2026
- 3D Software Synthesis Driven by Constraint-Expressive Intermediate RepresentationShuqing Li, Anson Y. Lam, Yun Peng, Wenxuan Wang et al.ICSE 2026
Builds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
- Divide-and-Conquer: Generating UI Code from ScreenshotsYuxuan Wan, Chaozheng Wang, Yi Dong, Wenxuan Wang et al.FSE 2025 · 10 citations
- CogAgent: A Visual Language Model for GUI AgentsWenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu et al.CVPR 2024
Related papers
- Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Fine-grained FeedbackJingyu Xiao, Jiantong Qin, Shuoqi Li, Man Ho Lam et al.KDD 2026 · 1 citation
- LaTCoder: Converting Webpage Design to Code with Layout-as-ThoughtYi Gui, Zhen Li, Zhongyi Zhang, Guohao Wang et al.KDD 2025 · 1 citation
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage DesignsYi Gui, Yao Wan, Zhen Li, Zhongyi Zhang et al.WWW 2025 · 24 citations
- Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive PrototypingJingyu Xiao, Yuxuan Wan, Yintong Huo, Zixin Wang et al.ASE 2025 · 1 citation
- Figma2Code: Automating Multimodal Design to Code in the WildYi Gui, Jiawan Zhang, Yina Wang, Tianran Ma et al.ICLR 2026 · 3 citations
