UI-Lens: Assessing General MLLMs' Potential to Automate UI Display Quality Assurance
Wei Xiang, Yexinrui Wu, Xinli Chen, Xinran Li, Shi Chen
Abstract
The goal of multimodal large language models (MLLMs) is not simply recognition, but visual discernment. This includes the critical capability to distinguish whether content is partially rendered or merely obscured by another element. Current models, which excel at object identification, fundamentally struggle with this ability, limiting their robustness in complex, layered environments. This challenge is acutely evident in User Interface (UI) display defect detection, which requires fine-grained element boundary understanding, missing-content detection, and reasoning about sequential interface semantic consistency. However, the capabilities of multimodal large language models (MLLMs) and vision-language models (VLMs) for detecting UI defect in realistic and complex interfaces have not been systematically validated. To fill this gap, we present UI-Lens, a UI display defect detection benchmark for multilingual UI scenarios. The dataset consists of 4,759 Chinese and 3,392 English interfaces meticulously annotated by design experts, covering six display defect categories. We conduct a systematic evaluation of 9 mainstream models (7 closedsource, 2 open-source). Results show clear shortcomings in current models: for tasks requiring fine-grained element boundary understanding, performance is near-random, with task-average F1 scores of 22.19% and 33.75% on Text Overflow and Container Overlap, respectively; for sequential interface semantic consistency, the task-average F1 score is only 11.44%, indicating severe underperformance. We release UI-Lens to catalyze research toward robust UI display defect detection with fine-grained boundary awareness in realistic interfaces. The dataset is available on this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe78bae2-72e4-4ad9-bc17-7e98c47f8f3fBuilds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Fine-Grained Visual PromptingLingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang et al.NeurIPS 2023 · 129 citations
- Owl Eyes: Spotting UI Display Issues via Visual UnderstandingZhe Liu, Chunyang Chen, Junjie Wang, Yuekai Huang et al.ASE 2020 · 79 citations
- A Theoretical Understanding of Self-Correction through In-context AlignmentYifei Wang, Yuyang Wu, Zeming Wei, Stefanie Jegelka et al.NeurIPS 2024 · 69 citations
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng et al.CHI 2023 · 49 citations
Related papers
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long et al.ICLR 2026 · 8 citations
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design UnderstandingJaehyun Jeon, Min Soo Kim, Janghan Yoon, Sumin Shim et al.ACL 2026 · 1 citation
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code GenerationRabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li et al.EMNLP 2025
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI AgentsRuihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong et al.ACL 2026 · 3 citations
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li et al.AAAI 2026 · 2 citations
