Automated Classification, Root Cause Analysis, and Repair Recommendations for Failed Mobile Testing by Specialized LLM
Chun Li, Fei Wang, Minxue Pan, Zhong Li, Mengliang Zeng, Bin Zhang, Xuejiao Yu, Boyun Wang, Kaijian Hua, Xuandong Li
摘要
Engineers utilize test scripts to conduct testing on mobile applications to ensure software reliability. In industrial settings, failures in mobile application testing stem not only from faults in the program or scripts but also from other aspects, such as the testing environment, making it challenging to apply existing automated analysis methods in such settings. Furthermore, mobile testing in industrial contexts generates a large volume of artifacts, such as screenshots and logs, and involves diverse device models, varied test environments, and complex functionalities. All these factors make the manual analysis of failed mobile tests a process that is typically time-consuming, costly, and labor-intensive. Large language models (LLMs) with strong logical reasoning capabilities offer a novel perspective for automated analysis of failed mobile tests. However, due to their limited knowledge of mobile testing, general LLMs can only achieve sub-optimal effectiveness in these tasks. To address these challenges, we propose AnaDroid, a novel LLM training framework specifically tailored for analyzing failed mobile tests. AnaDroid performs feature preprocessing for failed mobile tests and trains the large language model through two novel, tailored training procedures. Specifically, AnaDroid first leverages bidirectional pre-training to instill the domain-specific knowledge and reasoning capabilities required for failed test analysis into the model. AnaDroid further employs preference optimization to enhance the model's capacity to capture critical information within the input context, thereby further improving its analytical effectiveness. Through extensive evaluations across six open-source LLMs of diverse architectures and scales on a large-scale industrial dataset of 284k failed mobile testing scripts, AnaDroid outperformed all four baselines across three analysis tasks. Furthermore, during a month-long deployment within a global company’s testing system, AnaDroid demonstrated its practical utility by successfully providing root cause and repair recommendations for 78% and 81% of the 11,304 failed tests, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Generating Failure-Based Oracles to Support Testing of Reported Bugs in Android AppsJack Johnson, Junayed Mahmud, Oscar Chaparro, Kevin Moran 等ASE 2025 · 被引用 1 次
- OptiMine: Scalable and Precise Code Optimization for Android Apps via LLM-Driven Semantic AnalysisPengbo Du, Qiuping Yi, Liangzheng Zhang, Hongliang LiangISSTA 2026
- LLMDroid: Enhancing Automated Mobile App GUI Testing Coverage with Large Language Model GuidanceChenxu Wang, Tianming Liu, Yanjie Zhao, Minghui Yang 等FSE 2025 · 被引用 6 次
- Automated Program Repair for UI-Centric Android Bugs: How Far Are We?Junayed Mahmud, Sparsh Pandey, Nadeeshan De Silva, Atish Kumar Dipongkor 等ISSTA 2026
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等ICSE 2024 · 被引用 81 次
