Deep GUI: Black-box GUI Input Generation with Deep Learning
Faraz Yazdani Banafshe Daragh, Sam Malek
Abstract
Despite the proliferation of Android testing tools, Google Monkey has remained the de facto standard for practitioners. The popularity of Google Monkey is largely due to the fact that it is a black-box testing tool, making it widely applicable to all types of Android apps, regardless of their underlying implementation details. An important drawback of Google Monkey, however, is the fact that it uses the most naive form of test input generation technique, i.e., random testing. In this work, we present Deep GUI, an approach that aims to complement the benefits of black-box testing with a more intelligent form of GUI input generation. Given only screenshots of apps, Deep GUI first employs deep learning to construct a model of valid GUI interactions. It then uses this model to generate effective inputs for an app under test without the need to probe its implementation details. Moreover, since the data collection, training, and inference processes are performed independent of the platform, the model inferred by Deep GUI has application for testing apps in other platforms as well. We implemented a prototype of Deep GUI in a tool called Monkey++ by extending Google Monkey and evaluated it for its ability to crawl Android apps. We found that Monkey++ achieves significant improvements over Google Monkey in cases where an app’s UI is complex, requiring sophisticated inputs. Furthermore, our experimental results demonstrate the model inferred using Deep GUI can be reused for effective GUI input generation across platforms without the need for retraining.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60d826b7-da2c-4320-a9b6-6770979c0a39Cited by top-tier papers4
- Avgust: automating usage-based test generation from videos of app executionsYixue Zhao, Saghar Talebipour, Kesina Baral, Hyojae Park et al.FSE 2022 · 16 citations
- LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsShanhui Zhao, Hao Wen, Wenjie Du, Cheng Liang et al.MobiCom 2025 · 6 citations
- Navigating Mobile Testing Evaluation: A Comprehensive Statistical Analysis of Android GUI Testing MetricsYuanhong Lan, Yifei Lu, Minxue Pan, Xuandong LiASE 2024 · 4 citations
- PlayCoder: Making LLM-Generated GUI Code PlayableZhiyuan Peng, Wei Tao, Xin Yin, Chenhao Ying et al.FSE 2026
Builds on2
Related papers
- GUIFuzz++: Unleashing Grey-box Fuzzing on Desktop Graphical User Interfacing ApplicationsDillon Otto, Tanner Rowlett, Stefan NagyASE 2025
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- Seven Reasons Why: An In-Depth Study of the Limitations of Random Test Input Generation for AndroidFarnaz Behrang, Alessandro OrsoASE 2020 · 11 citations
- Deeply Reinforcing Android GUI Testing with Deep Reinforcement LearningYuanhong Lan, Yifei Lu, Zhong Li, Minxue Pan et al.ICSE 2024 · 21 citations
- Efficiency Matters: Speeding Up Automated Testing with GUI Rendering InferenceSidong Feng, Mulong Xie, Chunyang ChenICSE 2023 · 23 citations
