Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
Houston H. Zhang, Tao Zhang, Baoze Lin, Yuanqi Xue, Yincheng Zhu, Huan Liu, Li Gu, Linfeng Ye, Ziqiang Wang, Xinxin Zuo, Yang Wang, Yuanhao Yu, Zhixiang Chi
摘要
User interface to code (UI2Code) aims to generate executable code that can faithfully reconstruct a given input UI. Prior work focuses largely on web pages and mobile screens, leaving app widgets underexplored. Unlike web or mobile UIs with rich hierarchical context, widgets are compact, context-free micro-interfaces that summarize key information through dense layouts and iconography under strict spatial constraints. Moreover, while (image, code) pairs are widely available for web or mobile UIs, widget designs are proprietary and lack accessible markup. We formalize this setting as the Widget-to-Code (Widget2Code) and introduce an image-only widget benchmark with fine-grained, multi-dimensional evaluation metrics. Benchmarking shows that although generalized multimodal large language models (MLLMs) outperform specialized UI2Code methods, they still produce unreliable and visually inconsistent code. To address these limitations, we develop a baseline that jointly advances perceptual understanding and structured code generation. At the perceptual level, we follow widget design principles to assemble atomic components into complete layouts, equipped with icon retrieval and reusable visualization modules. At the system level, we design an end-to-end infrastructure, WidgetFactory, which includes a framework-agnostic widget-tailored domain-specific language (WidgetDSL) and a compiler that translates it into multiple front-end implementations (e.g., React, HTML/CSS). An adaptive rendering module further refines spatial dimensions to satisfy compactness constraints. Together, these contributions substantially enhance visual fidelity, establishing a strong baseline and unified infrastructure for future Widget2Code research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ASMIL: Attention-Stabilized Multiple Instance Learning for Whole-Slide ImagingLinfeng Ye, Shayan Mohajer Hamidi, Zhixiang Chi, Guang Li 等ICLR 2026 · 被引用 9 次
- MulFCoder: Framework-conditioned Multi-agent for MLLM-based Multi-framework Front-end Code GenerationJie Wu, Haoran Ma, Shisong Tang, Yulin Xu 等ICML 2026
- CL-DPS: A Contrastive Learning Approach to Blind Nonlinear Inverse Problem Solving via Diffusion Posterior SamplingLinfeng Ye, Shayan Mohajer Hamidi, Mert Pilanci, Konstantinos N. PlataniotisICLR 2026
它引用的顶会 Paper23
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid InferenceZhihang Lin, Mingbao Lin, Luxi Lin, Rongrong JiAAAI 2025 · 被引用 121 次
- MetaGCD: Learning to Continually Learn in Generalized Category DiscoveryYanan Wu, Zhixiang Chi, Yang Wang, Songhe FengICCV 2023 · 被引用 50 次
- Test-Time Domain Adaptation by Learning Domain-Aware Batch NormalizationYanan Wu, Zhixiang Chi, Yang Wang, Konstantinos N. Plataniotis 等AAAI 2024 · 被引用 41 次
相关 Paper
- MLLM-Based UI2Code Automation Guided by UI Layout InformationFan Wu, Cuiyun Gao, Shuqing Li, Xin-Cheng Wen 等ISSTA 2025 · 被引用 5 次
- Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Fine-grained FeedbackJingyu Xiao, Jiantong Qin, Shuoqi Li, Man Ho Lam 等KDD 2026 · 被引用 1 次
- Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code GenerationJiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang 等ACL 2026 · 被引用 2 次
- Figma2Code: Automating Multimodal Design to Code in the WildYi Gui, Jiawan Zhang, Yina Wang, Tianran Ma 等ICLR 2026 · 被引用 3 次
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage DesignsYi Gui, Yao Wan, Zhen Li, Zhongyi Zhang 等WWW 2025 · 被引用 24 次
