Lune

ISSTA2026Top-tier venue

On the Evaluation of Large Language Models in Unit Test Evolution (Experience Paper)

Weichang Liu, Junwei Zhang, Yuqing Niu, Bo Zhou

2026Year

Abstract

Large language models (LLMs) have recently shown promising potential in automating unit test evolution for evolving software systems. However, the effectiveness of LLMs in unit test evolution remains insufficiently understood, particularly with respect to prompt design choices, in-context learning (ICL) strategies, and different types of test evolution. In this paper, we present the first comprehensive empirical study to evaluate LLMs for unit test evolution. We systematically assess nine open-source code LLMs (3B to 34B parameters) and three state-of-the-art commercial models across diverse prompt designs, ICL strategies, and representative test evolution frameworks. To support robust and execution-based evaluation, we construct a new benchmark consisting of 530 real-world focal method–test co-evolution instances collected from seven actively maintained open-source projects. Our evaluation employs a suite of compilation, execution, and coverage-based metrics. Extensive experimental results reveal that prompt design and ICL methods significantly impact LLM effectiveness. Furthermore, the optimal configurations of these strategies vary substantially across different LLMs and evolution types. Based on our findings, we derive actionable insights to guide future research and practical adoption of LLM-based techniques for unit test evolution.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines