Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware Decisions
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, Qing Wang
Abstract
Automated Graphical User Interface (GUI) testing plays a crucial role in ensuring app quality, especially as mobile applications have become an integral part of our daily lives. Despite the growing popularity of learning-based techniques in automated GUI testing due to their ability to generate human-like interactions, they still suffer from several limitations, such as low testing coverage, inadequate generalization capabilities, and heavy reliance on training data. Inspired by the success of Large Language Models (LLMs) like ChatGPT in natural language understanding and question answering, we formulate the mobile GUI testing problem as a Q&A task. We propose GPTDroid, asking LLM to chat with the mobile apps by passing the GUI page information to LLM to elicit testing scripts, and executing them to keep passing the app feedback to LLM, iterating the whole process. Within this framework, we have also introduced a functionality-aware memory prompting mechanism that equips the LLM with the ability to retain testing knowledge of the whole process and conduct long-term, functionality-based reasoning to guide exploration. We evaluate it on 93 apps from Google Play and demonstrate that it outperforms the best baseline by 32% in activity coverage, and detects 31% more bugs at a faster rate. Moreover, GPTDroid identifies 53 new bugs on Google Play, of which 35 have been confirmed and fixed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88226a0b-ae16-4fad-be63-c10257faf55fCited by top-tier papers34
- Demystifying LLM-Based Software Engineering AgentsChunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming ZhangFSE 2025 · 36 citations
- From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered AnalysisZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ziang Xiao et al.CHI 2025 · 30 citations
- Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLMZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.CHI 2024 · 29 citations
- Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News DetectionChaowei Zhang, Zongling Feng, Zewei Zhang, Jipeng Qiang et al.AAAI 2025 · 13 citations
- LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsShanhui Zhao, Hao Wen, Wenjie Du, Cheng Liang et al.MobiCom 2025 · 6 citations
Builds on23
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng et al.WWW 2022 · 488 citations
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 223 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- Reinforcement learning based curiosity-driven testing of Android applicationsMinxue Pan, An Huang, Guoxin Wang, Tian Zhang et al.ISSTA 2020 · 166 citations
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
Related papers
- LLMDroid: Enhancing Automated Mobile App GUI Testing Coverage with Large Language Model GuidanceChenxu Wang, Tianming Liu, Yanjie Zhao, Minghui Yang et al.FSE 2025 · 6 citations
- Beyond Static GUI Agent: Evolving LLM-based GUI Testing via Dynamic MemoryMengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang et al.ASE 2025 · 1 citation
- Think Outside the Box: Automating Inter-App Functionality Testing via Memory Implanting and ReasoningMengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang et al.ICSE 2026
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che et al.ICSE 2023 · 107 citations
- Guardian: A Runtime Framework for LLM-Based UI ExplorationDezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu et al.ISSTA 2024 · 13 citations
