On the Evaluation of Large Language Models in Unit Test Generation
Lin Yang, Chen Yang, Shutao Gao, Weijing Wang, Bo Wang, Qihao Zhu, Xiao Chu, Jianyi Zhou, Guangtai Liang, Qianxiang Wang, Junjie Chen
摘要
Unit testing is an essential activity in software development for verifying the correctness of software components. However, manually writing unit tests is challenging and time-consuming. The emergence of Large Language Models (LLMs) offers a new direction for automating unit test generation. Existing research primarily focuses on closed-source LLMs (e.g., ChatGPT and CodeX) with fixed prompting strategies, leaving the capabilities of advanced open-source LLMs with various prompting settings unexplored. Particularly, open-source LLMs offer advantages in data privacy protection and have demonstrated superior performance in some tasks. Moreover, effective prompting is crucial for maximizing LLMs' capabilities. In this paper, we conduct the first empirical study to fill this gap, based on 17 Java projects, five widely-used open-source LLMs with different structures and parameter sizes, and comprehensive evaluation metrics. Our findings highlight the significant influence of various prompt factors, show the performance of open-source LLMs compared to the commercial GPT-4 and the traditional Evosuite, and identify limitations in LLM-based unit test generation. We then derive a series of implications from our study to guide future research and practical use of LLM-based unit test generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- LLM Test Generation via Iterative Hybrid Program AnalysisSijia Gu, Noor Nashid, Ali MesbahICSE 2026 · 被引用 6 次
- Retrieval-Augmented Test Generation: How Far Are We?Jiho Shin, Nima Shiri Harzevili, Reem Aleithan, Hadi Hemmati 等ICSE 2026 · 被引用 5 次
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li 等ICLR 2026 · 被引用 4 次
- Are Autonomous Web Agents Good Testers?Antoine Chevrot, Alexandre Vernotte, Jean-Rémy Falleri, Xavier Blanc 等ISSTA 2025 · 被引用 3 次
- QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code TranslationChangxin Ke, Rui Zhang, Shuo Wang, Li Ding 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- On the Evaluation of Large Language Models in Unit Test Evolution (Experience Paper)Weichang Liu, Junwei Zhang, Yuqing Niu, Bo ZhouISSTA 2026
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang 等FSE 2024 · 被引用 89 次
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu 等ISSTA 2025 · 被引用 7 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
- Test Intention Guided LLM-Based Unit Test GenerationZifan Nan, Zhaoqiang Guo, Kui Liu, Xin XiaICSE 2025 · 被引用 5 次
