Lune

ICML2026Top-tier venue

VisionWebDev: A Hierarchical Benchmark for Visual Website Development with Agent Verification

Zehai He, Wenyi Hong, ZHEN YANG, Ziyang Pan, Mingdao Liu, Xiaotao Gu, Jie Tang

2026Year
9Citations

Abstract

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical benchmark for visual website development, spanning from static UI-to-code generation, interactive multi-page frontend reproduction, to long-horizon full-stack website development. The benchmark is constructed from real-world websites and comprises a total of 193 tasks across 16 categories, with 918 prototype images and 1,255 test cases. To support flexible, thorough and reliable evaluation, we propose a workflow-based agent verification paradigm based on two complementary components: a GUI agent verifier and a VLM-based judge. We evaluate multiple visual language models instantiated under different coding-agent frameworks, revealing substantial performance gaps at all task levels, with state-of-the-art models still struggling on full-stack development. Project page: https: //vision2web-bench.github.io/ .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 21a6803a-261c-44d7-9f21-482fba477a95

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines