WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
Chenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang, Tao Xie
Abstract
Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential. However, building a benchmark for LLM-generated web apps remains challenging due to the need for real-world user requirements, generalizable evaluation metrics without relying on groundtruth implementations or test cases, and interpretable evaluation results. To address these challenges, we introduce WebCoderBench, the first real-world-collected, generalizable, and interpretable benchmark for web app generation. WebCoderBench comprises 1,572 real user requirements, covering diverse modalities and expression styles that reflect realistic user intentions. WebCoderBench provides 24 fine-grained evaluation metrics across 9 perspectives, combining rule-based and LLM-asa-judge paradigm for fully automated, objective, and general evaluation. Moreover, Web-CoderBench adopts human-preference-aligned weights over metrics to yield interpretable overall scores. Experiments across 12 representative LLMs and 2 LLM-based agents show that there exists no dominant model across all evaluation metrics, offering an opportunity for LLM developers to optimize their models in a targeted manner for a more powerful version. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00a41e68-b3db-4cb0-99a2-7541724ef8f6Builds on2
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong et al.ICLR 2026 · 150 citations
- Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal PredictionHuanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao et al.EMNLP 2025 · 1 citation
Related papers
- MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered AssistantsZuhao zhang, Chengyue Yu, Yuante Li, Chenyi Zhuang et al.ICML 2026 · 2 citations
- EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code EditsWayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal et al.ICLR 2026 · 7 citations
- Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation BenchmarkDewu Zheng, Yanlin Wang, Ensheng Shi, Xilin Liu et al.ICSE 2026
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 3 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
