On the Evaluation of Large Language Models in Unit Test Evolution (Experience Paper)
Weichang Liu, Junwei Zhang, Yuqing Niu, Bo Zhou
摘要
Large language models (LLMs) have recently shown promising potential in automating unit test evolution for evolving software systems. However, the effectiveness of LLMs in unit test evolution remains insufficiently understood, particularly with respect to prompt design choices, in-context learning (ICL) strategies, and different types of test evolution. In this paper, we present the first comprehensive empirical study to evaluate LLMs for unit test evolution. We systematically assess nine open-source code LLMs (3B to 34B parameters) and three state-of-the-art commercial models across diverse prompt designs, ICL strategies, and representative test evolution frameworks. To support robust and execution-based evaluation, we construct a new benchmark consisting of 530 real-world focal method–test co-evolution instances collected from seven actively maintained open-source projects. Our evaluation employs a suite of compilation, execution, and coverage-based metrics. Extensive experimental results reveal that prompt design and ICL methods significantly impact LLM effectiveness. Furthermore, the optimal configurations of these strategies vary substantially across different LLMs and evolution types. Based on our findings, we derive actionable insights to guide future research and practical adoption of LLM-based techniques for unit test evolution.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang 等ASE 2024 · 被引用 42 次
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu 等ISSTA 2025 · 被引用 7 次
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du 等ICSE 2026
- Test Intention Guided LLM-Based Unit Test GenerationZifan Nan, Zhaoqiang Guo, Kui Liu, Xin XiaICSE 2025 · 被引用 5 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
