Feedback-Driven Automated Whole Bug Report Reproduction for Android Apps
Dingbang Wang, Yu Zhao, Sidong Feng, Zhaoxu Zhang, William G. J. Halfond, Chunyang Chen, Xiaoxia Sun, Jiangfan Shi, Tingting Yu
Abstract
In software development, bug report reproduction is a challenging task. This paper introduces ReBL, a novel feedback-driven approach that leverages GPT-4, a large-scale language model (LLM), to automatically reproduce Android bug reports. Unlike traditional methods, ReBL bypasses the use of Step to Reproduce (S2R) entities. Instead, it leverages the entire textual bug report and employs innovative prompts to enhance GPT's contextual reasoning. This approach is more flexible and context-aware than the traditional step-by-step entity matching approach, resulting in improved accuracy and effectiveness. In addition to handling crash reports, ReBL has the capability of handling non-crash functional bug reports. Our evaluation of 96 Android bug reports (73 crash and 23 non-crash) demonstrates that ReBL successfully reproduced 90.63% of these reports, averaging only 74.98 seconds per bug report. Additionally, ReBL outperformed three existing tools in both success rate and speed. CCS Concepts • Software and its engineering → Software testing and debugging;
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e0463c6-120d-4663-b7d6-acdf630268eaCited by top-tier papers9
- Agents in the Sandbox: End-to-End Crash Bug Reproduction for MinecraftEray Yapagci, Yavuz Alp Sencer Öztürk, Eray TüzünASE 2025 · 3 citations
- Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature TestingSidong Feng, Changhao Du, Huaxiao Liu, Qingnan Wang et al.ICSE 2026 · 2 citations
- Automated Test Transfer across Android Apps using Large Language ModelsBenyamin Beyzaei, Saghar Talebipour, Ghazal Rafiei, Nenad Medvidovic et al.ISSTA 2025 · 2 citations
- Towards Automated Crowdsourced Testing via Personified-LLMShengcheng Yu, Yuchen Ling, Chunrong Fang, Zhenyu Chen et al.FSE 2026 · 1 citation
- Generating Failure-Based Oracles to Support Testing of Reported Bugs in Android AppsJack Johnson, Junayed Mahmud, Oscar Chaparro, Kevin Moran et al.ASE 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 143 citations
- Translating video recordings of mobile app usages into replayable scenariosCarlos Bernal-Cárdenas, Nathan Cooper, Kevin Moran, Oscar Chaparro et al.ICSE 2020 · 61 citations
Related papers
- Automatically Reproducing Android Bug Reports using Natural Language Processing and Reinforcement LearningZhaoxu Zhang, Robert Winn, Yu Zhao, Tingting Yu et al.ISSTA 2023 · 14 citations
- CrashTranslator: Automatically Reproducing Mobile Application Crashes Directly from Stack TraceYuchao Huang, Junjie Wang, Zhe Liu, Yawen Wang et al.ICSE 2024 · 22 citations
- Context-aware Bug Reproduction for Mobile AppsYuchao Huang, Junjie Wang, Zhe Liu, Song Wang et al.ICSE 2023 · 7 citations
- Enhancing LLM-Based Bug Reproduction via Code Entity Retrieval and Test Case RepairHao Ding, Yanjie Jiang, Yuxia Zhang, Hui LiuISSTA 2026
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent AgentMehil Shah, Mohammad Masudur Rahman, Foutse KhomhICSE 2026
