WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements
Xiwen Teoh, Yun Lin, Duc-Minh Nguyen, Ruofei Ren, Wenjie Zhang, Jin Song Dong
摘要
Visual language model (VLM) agents show great promise in automating graphical user interface (GUI) testing against requirements in natural language. However, the probabilistic nature of language models can have inherent hallucinations. Therefore, given a detected inconsistency between the requirement and the web application, it is hard to distinguish whether it stems from the hallucination or a real application bug. Addressing this issue presents two core technical challenges: (1) limited capability and accuracy in deriving implicit test oracles, where the agent must act as its own oracle to implicitly decide if the application's behavior is correct without guidance, and (2) limited reliability due to probabilistic inference, where an LLM's inconsistent reasoning undermines its trustworthiness as an oracle.
We introduce WebTestPilot, a neurosymbolic LLM-based approach that addresses both challenges through symbolization. WebTestPilot detects and abstracts critical GUI elements of a web application into symbolic variables. This design improves reliability by constraining assertion generation to operations grounded in explicitly defined symbols, thereby reducing unconstrained or inconsistent reasoning. At the same time, it improves accuracy by representing application states and their relationships in a structured symbolic form, which increases the likelihood of the agent recognizing data, causal, and temporal dependencies across states. Together, these capabilities enable WebTestPilot to generate reliable and accurate test oracles that capture meaningful implicit expectations derived from test requirements. To advance research in this area, we build a benchmark of bug-injected web apps for evaluating NL-to-E2E testing. The results show that WebTestPilot achieves a task completion rate of 99%, with 96% precision and 96% recall in bug detection, outperforming the best baseline (+70 precision, +27 recall). The agent generalizes across diverse natural language inputs (i.e., those containing typos, grammatical errors, redundant sentences, stylistic restyling, or abbreviations) and model scales (3B-72B). In a real-world deployment with a no-code platform, WebTestPilot discovered 8 bugs during development, including data binding, UI, and navigation issues.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
- Time-travel testing of Android appsZhen Dong, Marcel Böhme, Lucia Cojocaru, Abhik RoychoudhuryICSE 2020 · 被引用 104 次
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等ICSE 2024 · 被引用 81 次
- Benchmarking automated GUI testing for Android against real-world bugsTing Su, Jue Wang, Zhendong SuFSE 2021 · 被引用 77 次
- Automatic Web Testing Using Curiosity-Driven Reinforcement LearningYan Zheng, Yi Liu, Xiaofei Xie, Yepang Liu 等ICSE 2021 · 被引用 75 次
相关 Paper
- GUIPilot: A Consistency-Based Mobile GUI Testing Approach for Detecting Application-Specific BugsRuofan Liu, Xiwen Teoh, Yun Lin, Guanjie Chen 等ISSTA 2025 · 被引用 5 次
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage DesignsYi Gui, Yao Wan, Zhen Li, Zhongyi Zhang 等WWW 2025 · 被引用 24 次
- WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic ExplorationYao Zhang, Zijian Ma, Yunpu Ma, Zhen Han 等AAAI 2025 · 被引用 101 次
- AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsHongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen 等ACL 2025
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
