UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
ZHEN YANG, Wenyi Hong, Mingde Xu, Xinyue Fan, Weihan Wang, Jiale Cheng, Xiaotao Gu, Jie Tang
Abstract
UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing methods formulate UI-to-code as a single-pass generation, which mismatches real-world UI development that is inherently iterative and feedback-driven. We reformulate UI-to-code as an interactive visual optimization problem, where code generation is embedded in a closed-loop process of execution, visual inspection, and iterative refinement driven by rendered visual feedback. To address the non-differentiability of visual objectives and the noise of absolute visual evaluators, we propose Relative Visual Policy Optimization (RVPO), a preference-based reinforcement learning method that optimizes relative visual rankings among rendered candidates under execution feedback. We instantiate this paradigm in UI2Code, an open-source 9B model trained via continual pre-training, supervised fine-tuning, and reinforcement learning. Experiments demonstrate state-of-the-art performance on UI drafting, UI polishing, and UI editing benchmarks, even outperforming larger models, with performance consistently improving through iterative visual optimization. Our code and models are available at https://github.com/zai-org/UI2Code_N.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a3a6275-8d12-4779-86d3-bb5af753dd06Cited by top-tier papers3
- Seeing Is Coding: On the Effectiveness of Vision Language Models in Code UnderstandingYuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen et al.ISSTA 2026 · 1 citation
- MulFCoder: Framework-conditioned Multi-agent for MLLM-based Multi-framework Front-end Code GenerationJie Wu, Haoran Ma, Shisong Tang, Yulin Xu et al.ICML 2026
- Deterministic Component Mining for Multi-Framework UI2Code GenerationZixiong Yang, Linxiao Li, Jiaye Lin, Binrui Wu et al.ICML 2026
Builds on6
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form PreferencesZhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li et al.ICLR 2026 · 11 citations
- Divide-and-Conquer: Generating UI Code from ScreenshotsYuxuan Wan, Chaozheng Wang, Yi Dong, Wenxuan Wang et al.FSE 2025 · 10 citations
Related papers
- Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided RefinementZhihan Zhang, Yixin Cao, Lizi LiaoACM MM 2025
- Seeing is Improving: Visual Feedback for Iterative Text Layout RefinementJunrong Guo, Shancheng Fang, Yadong Qu, Hongtao XieCVPR 2026 · 2 citations
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy OptimizationXinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li et al.CVPR 2026 · 25 citations
- On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language UnderstandingHaoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao et al.NeurIPS 2025 · 2 citations
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationMingzhe Du, Anh Tuan Luu, Yue Liu, Yuhao Qing et al.NeurIPS 2025 · 18 citations
