Guardian: A Runtime Framework for LLM-Based UI Exploration
Dezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu, Yuan Cao, Ying Zhang, Wei Yang, Tao Xie
摘要
Tests for feature-based UI testing have been indispensable for ensuring the quality of mobile applications (apps for short). The high manual labor costs to create such tests have led to a strong interest in automated feature-based UI testing, where an approach automatically explores the App under Test (AUT) to find correct sequences of UI events achieving the target test objective, given only a high-level test objective description. Given that the task of automated feature-based UI testing resembles conventional AI planning problems, large language models (LLMs), known for their effectiveness in AI planning, could be ideal for this task. However, our study reveals that LLMs struggle with following specific instructions for UI testing and replanning based on new information. This limitation results in reduced effectiveness of LLM-driven solutions for automated feature-based UI testing, despite the use of advanced prompting techniques. Toward addressing the preceding limitation, we propose Guardian, a runtime system framework to improve the effectiveness of automated feature-based UI testing by offloading computational tasks from LLMs with two major strategies. First, Guardian refines UI action space that the LLM can plan over, enforcing the instruction following of the LLM by construction. Second, Guardian deliberately checks whether the gradually enriched information invalidates previous planning by the LLM. Guardian removes the invalidated UI actions from the UI action space that the LLM can plan over, restores the state of the AUT to the state before the execution of the invalidated UI actions, and prompts the LLM to re-plan with the new UI action space. We instantiate Guardian with ChatGPT and construct a benchmark named FestiVal with 58 tasks from 23 highly popular apps. Evaluation results on FestiVal show that Guardian achieves 48.3
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- From Assistant to Independent Developer — Are GPTs Ready for Software Development?Dezhi Ran, Yuan Cao, Mengzhou Wu, Simin Chen 等ICLR 2026 · 被引用 4 次
- Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature TestingSidong Feng, Changhao Du, Huaxiao Liu, Qingnan Wang 等ICSE 2026 · 被引用 2 次
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
- TaOPT: Tool-Agnostic Optimization of Parallelized Automated Mobile UI TestingDezhi Ran, Zihe Song, Wenyu Wang, Wei Yang 等ASPLOS 2025 · 被引用 2 次
- Automating UI Optimization through Multi-Agentic ReasoningZhipeng Li, Christoph Gebhardt, Yi-Chi Liao, Christian HolzCHI 2026 · 被引用 1 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
相关 Paper
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等ICSE 2024 · 被引用 81 次
- LLMDroid: Enhancing Automated Mobile App GUI Testing Coverage with Large Language Model GuidanceChenxu Wang, Tianming Liu, Yanjie Zhao, Minghui Yang 等FSE 2025 · 被引用 6 次
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang 等ASE 2024 · 被引用 42 次
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 等EMNLP 2024 · 被引用 6 次
