UI-Lens: Assessing General MLLMs' Potential to Automate UI Display Quality Assurance
Wei Xiang, Yexinrui Wu, Xinli Chen, Xinran Li, Shi Chen
摘要
The goal of multimodal large language models (MLLMs) is not simply recognition, but visual discernment. This includes the critical capability to distinguish whether content is partially rendered or merely obscured by another element. Current models, which excel at object identification, fundamentally struggle with this ability, limiting their robustness in complex, layered environments. This challenge is acutely evident in User Interface (UI) display defect detection, which requires fine-grained element boundary understanding, missing-content detection, and reasoning about sequential interface semantic consistency. However, the capabilities of multimodal large language models (MLLMs) and vision-language models (VLMs) for detecting UI defect in realistic and complex interfaces have not been systematically validated. To fill this gap, we present UI-Lens, a UI display defect detection benchmark for multilingual UI scenarios. The dataset consists of 4,759 Chinese and 3,392 English interfaces meticulously annotated by design experts, covering six display defect categories. We conduct a systematic evaluation of 9 mainstream models (7 closedsource, 2 open-source). Results show clear shortcomings in current models: for tasks requiring fine-grained element boundary understanding, performance is near-random, with task-average F1 scores of 22.19% and 33.75% on Text Overflow and Container Overlap, respectively; for sequential interface semantic consistency, the task-average F1 score is only 11.44%, indicating severe underperformance. We release UI-Lens to catalyze research toward robust UI display defect detection with fine-grained boundary awareness in realistic interfaces. The dataset is available on this link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Fine-Grained Visual PromptingLingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang 等NeurIPS 2023 · 被引用 129 次
- Owl Eyes: Spotting UI Display Issues via Visual UnderstandingZhe Liu, Chunyang Chen, Junjie Wang, Yuekai Huang 等ASE 2020 · 被引用 79 次
- A Theoretical Understanding of Self-Correction through In-context AlignmentYifei Wang, Yuyang Wu, Zeming Wei, Stefanie Jegelka 等NeurIPS 2024 · 被引用 69 次
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng 等CHI 2023 · 被引用 49 次
相关 Paper
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long 等ICLR 2026 · 被引用 8 次
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design UnderstandingJaehyun Jeon, Min Soo Kim, Janghan Yoon, Sumin Shim 等ACL 2026 · 被引用 1 次
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code GenerationRabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li 等EMNLP 2025
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI AgentsRuihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong 等ACL 2026 · 被引用 3 次
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li 等AAAI 2026 · 被引用 2 次
