Benchmarking automated GUI testing for Android against real-world bugs
Ting Su, Jue Wang, Zhendong Su
摘要
For ensuring the reliability of Android apps, there has been tremendous, continuous progress on improving automated GUI testing in the past decade. Specifically, dozens of testing techniques and tools have been developed and demonstrated to be effective in detecting crash bugs and outperform their respective prior work in the number of detected crashes. However, an overarching question łHow effectively and thoroughly can these tools find crash bugs in practice?ž has not been well-explored, which requires a ground-truth benchmark with real-world bugs. Since prior studies focus on tool comparisons w.r.t. some selected apps, they cannot provide direct, in-depth answers to this question.
To complement existing work and tackle the above question, this paper offers the first ground-truth empirical evaluation of automated GUI testing for Android. To this end, we devote substantial manual effort to set up the Themis benchmark set, including (1) a carefully constructed dataset with 52 real, reproducible crash bugs (taking two person-months for its collection and validation), and (2) a unified, extensible infrastructure with six recent state-of-the-art testing tools. The whole evaluation has taken over 10,920 CPU hours. We find a considerable gap in these tools finding the collected real bugs Ð 18 bugs cannot be detected by any tool. Our systematic analysis further identifies five major common challenges that these tools face, and reveals additional findings such as factors affecting these tools in bug finding and opportunities for tool improvements. Overall, this work offers new concrete insights, most of which are previously unknown/unstated and difficult to obtain. Our study presents a new, complementary perspective from prior studies to understand and analyze the effectiveness of existing testing tools, as well as a benchmark for future research on this topic. The Themis benchmark is publicly available at https:// github.com/ the-themis-benchmarks/ home.
• Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 被引用 143 次
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等ICSE 2024 · 被引用 81 次
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language ModelsChenyuan Yang, Yinlin Deng, Runyu Lu, Jiayi Yao 等OOPSLA 2024 · 被引用 74 次
- Fully automated functional fuzzing of Android apps for detecting non-crashing logic bugsTing Su, Yichen Yan, Jue Wang, Jingling Sun 等OOPSLA 2021 · 被引用 58 次
它引用的顶会 Paper7
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- LAVA: Large-Scale Automated Vulnerability AdditionBrendan Dolan-Gavitt, Patrick Hulin, Engin Kirda, Tim Leek 等S&P 2016 · 被引用 354 次
- Reinforcement learning based curiosity-driven testing of Android applicationsMinxue Pan, An Huang, Guoxin Wang, Tian Zhang 等ISSTA 2020 · 被引用 166 次
- Time-travel testing of Android appsZhen Dong, Marcel Böhme, Lucia Cojocaru, Abhik RoychoudhuryICSE 2020 · 被引用 104 次
- ComboDroid: generating high-quality test inputs for Android apps via use case combinationsJue Wang, Yanyan Jiang, Chang Xu, Chun Cao 等ICSE 2020 · 被引用 61 次
相关 Paper
- Automata-Based Trace Analysis for Aiding Diagnosing GUI Testing Tools for AndroidEnze Ma, Shan Huang, Weigang He, Ting Su 等FSE 2023 · 被引用 3 次
- General and Practical Property-based Testing for Android AppsYiheng Xiong, Ting Su, Jue Wang, Jingling Sun 等ASE 2024 · 被引用 5 次
- An Empirical Study of Functional Bugs in Android AppsYiheng Xiong, Mengqian Xu, Ting Su, Jingling Sun 等ISSTA 2023 · 被引用 40 次
- Mobile Application Coverage: The 30% Curse and Ways ForwardFaridah Akinotcho, Lili Wei, Julia RubinICSE 2025 · 被引用 3 次
- On Using GUI Interaction Data to Improve Text Retrieval-based Bug LocalizationJunayed Mahmud, Nadeeshan De Silva, Safwat Ali Khan, Seyed Hooman Mostafavi 等ICSE 2024 · 被引用 12 次
