Navigating Mobile Testing Evaluation: A Comprehensive Statistical Analysis of Android GUI Testing Metrics
Yuanhong Lan, Yifei Lu, Minxue Pan, Xuandong Li
摘要
The prominent role of mobile apps in daily life has underscored the need for robust quality assurance, leading to the development of various automated Android Graphical User Interface (GUI) testing approaches. Code coverage and fault detection are two primary metrics for evaluating the effectiveness of these testing approaches. However, conducting a reliable and robust evaluation based on the two metrics remains challenging, due to the imperfections of the current evaluation system, with a tangle of numerous metric granularities and the interference of multiple nondeterminism in tests. For instance, the evaluation solely based on the mean or total numbers of detected faults lacks statistical robustness, resulting in numerous conflicting conclusions that impede the comprehensive understanding of stakeholders involved in Android testing, thereby hindering the advancement of Android testing methodologies. To mitigate such issues, this paper presents the first comprehensive statistical study of existing Android GUI testing metrics, involving extensive experiments with 8 state-of-the-art testing approaches on 42 diverse apps, examining aspects including statistical significance, correlation, and variation. Our study focuses on two primary areas: (1) The statistical significance and correlation between test metrics and among different metric granularities. (2) The influence of test randomness and test convergence on evaluation results of test metrics. By employing statistical analysis to account for the considerable influence of randomness, we achieve notable findings: (1) Instruction, Executable Lines Of Code (ELOC), and method coverage demonstrate notable consistency across both significance evaluation and mean value evaluation, whereas the evaluation on Fatal Errors compared to Core Vitals, as well as all errors versus the well-selected errors, reveals a similarly high level of consistency. (2) There are evident inconsistencies in the code coverage and fault detection results, indicating both two metrics should be considered for comprehensive evaluation. (3) Code coverage typically exhibits greater stability and robustness in evaluation compared to fault detection, whereas fault detection is quite unstable even with the maximum test rounds ever used in previous research studies. (4) A moderate test duration is sufficient for most approaches to showcase their comprehensive overall effectiveness on most apps in both code coverage and fault detection, indicating the possibility of adopting a moderate test duration to draw preliminary conclusions in Android testing development. These findings inform practical recommendations and support our proposal of an effective framework to enhance future mobile testing evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VRExplorer: A Model-based Approach for Semi-Automated Testing of Virtual Reality ScenesZhengyang Zhu, Hong-Ning Dai, Hanyang Guo, Zeqin Liao 等ASE 2025 · 被引用 1 次
- Profile Coverage: Using Android Compilation Profiles to Evaluate Dynamic TestingJakob Bleier, Felix Kehrer, Jürgen Cito, Martina LindorferASE 2025
它引用的顶会 Paper8
- Reinforcement learning based curiosity-driven testing of Android applicationsMinxue Pan, An Huang, Guoxin Wang, Tian Zhang 等ISSTA 2020 · 被引用 166 次
- Time-travel testing of Android appsZhen Dong, Marcel Böhme, Lucia Cojocaru, Abhik RoychoudhuryICSE 2020 · 被引用 104 次
- Benchmarking automated GUI testing for Android against real-world bugsTing Su, Jue Wang, Zhendong SuFSE 2021 · 被引用 77 次
- ComboDroid: generating high-quality test inputs for Android apps via use case combinationsJue Wang, Yanyan Jiang, Chang Xu, Chun Cao 等ICSE 2020 · 被引用 61 次
- Deep GUI: Black-box GUI Input Generation with Deep LearningFaraz Yazdani Banafshe Daragh, Sam MalekASE 2021 · 被引用 29 次
相关 Paper
- Deeply Reinforcing Android GUI Testing with Deep Reinforcement LearningYuanhong Lan, Yifei Lu, Zhong Li, Minxue Pan 等ICSE 2024 · 被引用 21 次
- Measuring and Mitigating Gaps in Structural TestingSoneya Binta Hossain, Matthew B. Dwyer, Sebastian G. Elbaum, Anh Nguyen-TuongICSE 2023 · 被引用 7 次
- Mobile Application Coverage: The 30% Curse and Ways ForwardFaridah Akinotcho, Lili Wei, Julia RubinICSE 2025 · 被引用 3 次
- An infrastructure approach to improving effectiveness of Android UI testing toolsWenyu Wang, Wing Lam, Tao XieISSTA 2021 · 被引用 33 次
- Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set SizeYiqun T. Chen, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst 等ASE 2020 · 被引用 48 次
