Chatgpt-Based Test Generation for Refactoring Engines Enhanced by Feature Analysis on Examples
Chunhao Dong, Yanjie Jiang, Yuxia Zhang, Yang Zhang, Hui Liu
Abstract
Software refactoring is widely employed to improve software quality. However, conducting refactorings manually is tedious, time-consuming, and error-prone. Consequently, automated and semi-automated tool support is highly desirable for software refactoring in the industry, and most of the main-stream IDEs provide powerful tool support for refactoring. However, complex refactoring engines are prone to errors, which in turn may result in imperfect and incorrect refactorings. To this end, in this paper, we propose a ChatGPT-based approach to testing refactoring engines. We first manually analyze bug reports and test cases associated with refactoring engines, and construct a feature library containing fine-grained features that may trigger defects in refactoring engines. The approach automatically generates prompts according to both predefined prompt templates and features randomly selected from the feature library, requesting ChatGPT to generate test programs with the requested features. Test programs generated by ChatGPT are then forwarded to multiple refactoring engines for differential testing. To the best of our knowledge, it is the first approach in testing refactoring engines that guides test program generation with features derived from existing bugs. It is also the first approach in this line that exploits LLMs in the generation of test programs. Our initial evaluation of four main-stream refactoring engines suggests that the proposed approach is effective. It identified a total of 115 previously unknown bugs besides 28 inconsistent refactoring behaviors among different engines. Among the 115 bugs, 78 have been manually confirmed by the original developers of the tested engines, i.e., IntelliJ IDEA, Eclipse, VScode-Java, and NetBeans.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Generating Project-Specific Test Cases with Requirement Validation IntentionBinhang Qi, Yun Lin, Xinyi Weng, Yuhuan Huang et al.ISSTA 2026 · 3 citations
- Coverage-Based Harmfulness Testing for LLM Code TransformationHonghao Tan, Haibo Wang, Diany Pressato, Yisen Xu et al.ASE 2025 · 2 citations
- Generalizing Test Cases for Comprehensive Test Scenario CoverageBinhang Qi, Yun Lin, Xinyi Weng, Chenyan Liu et al.FSE 2026 · 1 citation
Related papers
- Cross-Refactoring-Type Test Program Migration for Refactoring EnginesChunhao Dong, Yanjie Jiang, Yang Zhang, Yuxia Zhang et al.FSE 2026
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software RefactoringKhouloud Oueslati, Maxime Lamothe, Foutse KhomhICSE 2026 · 1 citation
- Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential PromptingTsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian et al.ASE 2023 · 51 citations
- Chatgpt Inaccuracy Mitigation During Technical Report Understanding: Are we There Yet?Salma Begum Tamanna, Gias Uddin, Song Wang, Lan Xia et al.ICSE 2025 · 2 citations
