On the Evaluation of Large Language Models in Unit Test Evolution (Experience Paper)
Weichang Liu, Junwei Zhang, Yuqing Niu, Bo Zhou
Abstract
Large language models (LLMs) have recently shown promising potential in automating unit test evolution for evolving software systems. However, the effectiveness of LLMs in unit test evolution remains insufficiently understood, particularly with respect to prompt design choices, in-context learning (ICL) strategies, and different types of test evolution. In this paper, we present the first comprehensive empirical study to evaluate LLMs for unit test evolution. We systematically assess nine open-source code LLMs (3B to 34B parameters) and three state-of-the-art commercial models across diverse prompt designs, ICL strategies, and representative test evolution frameworks. To support robust and execution-based evaluation, we construct a new benchmark consisting of 530 real-world focal method–test co-evolution instances collected from seven actively maintained open-source projects. Our evaluation employs a suite of compilation, execution, and coverage-based metrics. Extensive experimental results reveal that prompt design and ICL methods significantly impact LLM effectiveness. Furthermore, the optimal configurations of these strategies vary substantially across different LLMs and evolution types. Based on our findings, we derive actionable insights to guide future research and practical adoption of LLM-based techniques for unit test evolution.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang et al.ASE 2024 · 42 citations
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu et al.ISSTA 2025 · 7 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Test Intention Guided LLM-Based Unit Test GenerationZifan Nan, Zhaoqiang Guo, Kui Liu, Xin XiaICSE 2025 · 5 citations
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
