Enhancing Exploratory Testing by Large Language Model and Knowledge Graph
Yanqi Su, Dianshu Liao, Zhenchang Xing, Qing Huang, Mulong Xie, Qinghua Lu, Xiwei Xu
Abstract
Exploratory testing leverages the tester's knowledge and creativity to design test cases for effectively uncovering system-level bugs from the end user's perspective. Researchers have worked on test scenario generation to support exploratory testing based on a system knowledge graph, enriched with scenario and oracle knowledge from bug reports. Nevertheless, the adoption of this approach is hindered by difficulties in handling bug reports of inconsistent quality and varied expression styles, along with the infeasibility of the generated test scenarios. To overcome these limitations, we utilize the superior natural language understanding (NLU) capabilities of Large Language Models (LLMs) to construct a System KG of User Tasks and Failures (SysKG-UTF). Leveraging the system and bug knowledge from the KG, along with the logical reasoning capabilities of LLMs, we generate test scenarios with high feasibility and coherence. Particularly, we design chain-of-thought (CoT) reasoning to extract human-like knowledge and logical reasoning from LLMs, simulating a developer's process of validating test scenario feasibility. Our evaluation shows that our approach significantly enhances the KG construction, particularly for bug reports with low quality. Furthermore, our approach generates test scenarios with high feasibility and coherence. The user study further proves the effectiveness of our generated test scenarios in supporting exploratory testing. Specifically, 8 participants find 36 bugs from 8 seed bugs in two hours using our test scenarios, a significant improvement over the 21 bugs found by the state-of-the-art baseline.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Automated Soap Opera Testing Directed by LLMs and Scenario Knowledge: Feasibility, Challenges, and Road AheadYanqi Su, Zhenchang Xing, Chong Wang, Chunyang Chen et al.FSE 2025 · 1 citation
- Automated Extraction and Analysis of Developer's Rationale in Open Source SoftwareMouna Dhaouadi, Bentley Oakes, Michalis FamelisFSE 2025
- RippleGUItester: Change-Aware Exploratory TestingYanqi Su, Michael Pradel, Chunyang ChenISSTA 2026
- E-Test: E'er-Improving Test SuitesKetai Qiu, Luca Di Grazia, Leonardo Mariani, Mauro PezzèICSE 2026
Related papers
- Constructing a System Knowledge Graph of User Tasks and Failures from Bug Reports to Support Soap Opera TestingYanqi Su, Zheming Han, Zhenchang Xing, Xin Xia et al.ASE 2022 · 8 citations
- Generating Failure-Based Oracles to Support Testing of Reported Bugs in Android AppsJack Johnson, Junayed Mahmud, Oscar Chaparro, Kevin Moran et al.ASE 2025 · 1 citation
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 143 citations
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test GenerationSteven Liu, Jane Luo, Xin Zhang, Aofan Liu et al.ICML 2026 · 4 citations
