Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware Decisions
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, Qing Wang
摘要
Automated Graphical User Interface (GUI) testing plays a crucial role in ensuring app quality, especially as mobile applications have become an integral part of our daily lives. Despite the growing popularity of learning-based techniques in automated GUI testing due to their ability to generate human-like interactions, they still suffer from several limitations, such as low testing coverage, inadequate generalization capabilities, and heavy reliance on training data. Inspired by the success of Large Language Models (LLMs) like ChatGPT in natural language understanding and question answering, we formulate the mobile GUI testing problem as a Q&A task. We propose GPTDroid, asking LLM to chat with the mobile apps by passing the GUI page information to LLM to elicit testing scripts, and executing them to keep passing the app feedback to LLM, iterating the whole process. Within this framework, we have also introduced a functionality-aware memory prompting mechanism that equips the LLM with the ability to retain testing knowledge of the whole process and conduct long-term, functionality-based reasoning to guide exploration. We evaluate it on 93 apps from Google Play and demonstrate that it outperforms the best baseline by 32% in activity coverage, and detects 31% more bugs at a faster rate. Moreover, GPTDroid identifies 53 new bugs on Google Play, of which 35 have been confirmed and fixed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Demystifying LLM-Based Software Engineering AgentsChunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming ZhangFSE 2025 · 被引用 36 次
- From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered AnalysisZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ziang Xiao 等CHI 2025 · 被引用 30 次
- Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLMZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等CHI 2024 · 被引用 29 次
- Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News DetectionChaowei Zhang, Zongling Feng, Zewei Zhang, Jipeng Qiang 等AAAI 2025 · 被引用 13 次
- LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsShanhui Zhao, Hao Wen, Wenjie Du, Cheng Liang 等MobiCom 2025 · 被引用 6 次
它引用的顶会 Paper23
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng 等WWW 2022 · 被引用 488 次
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 被引用 223 次
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- Reinforcement learning based curiosity-driven testing of Android applicationsMinxue Pan, An Huang, Guoxin Wang, Tian Zhang 等ISSTA 2020 · 被引用 166 次
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 被引用 163 次
相关 Paper
- LLMDroid: Enhancing Automated Mobile App GUI Testing Coverage with Large Language Model GuidanceChenxu Wang, Tianming Liu, Yanjie Zhao, Minghui Yang 等FSE 2025 · 被引用 6 次
- Beyond Static GUI Agent: Evolving LLM-based GUI Testing via Dynamic MemoryMengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang 等ASE 2025 · 被引用 1 次
- Think Outside the Box: Automating Inter-App Functionality Testing via Memory Implanting and ReasoningMengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang 等ICSE 2026
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
- Guardian: A Runtime Framework for LLM-Based UI ExplorationDezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu 等ISSTA 2024 · 被引用 13 次
