Domain Adaptation for Code Model-Based Unit Test Case Generation
Jiho Shin, Sepehr Hashtroudi, Hadi Hemmati, Song Wang
Abstract
Recently, deep learning-based test case generation approaches have been proposed to automate the generation of unit test cases. In this study, we leverage Transformer-based code models to generate unit tests with the help of Domain Adaptation (DA) at a project level. Specifically, we use CodeT5, a relatively small language model trained on source code data, and fine-tune it on the test generation task. Then, we apply domain adaptation to each target project data to learn project-specific knowledge (project-level DA). We use the Methods2test dataset to fine-tune CodeT5 for the test generation task and the Defects4j dataset for project-level domain adaptation and evaluation. We compare our approach with (a) CodeT5 finetuned on the test generation without DA, (b) the A3Test tool, and (c) GPT-4 on five projects from the Defects4j dataset. The results show that tests generated using DA can increase the line coverage by 18.62%, 19.88%, and 18.02% and mutation score by 16.45%, 16.01%, and 12.99% compared to the above (a), (b), and (c) baselines, respectively. The overall results show consistent improvements in metrics such as parse rate, compile rate, BLEU, and CodeBLEU. In addition, we show that our approach can be seen as a complementary solution alongside existing search-based test generation tools such as EvoSuite, to increase the overall coverage and mutation scores with an average of 34.42% and 6.8%, for line coverage and mutation score, respectively. CCS CONCEPTS • Software and its engineering → Software testing and debugging; Empirical software validation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4b28701-666e-4470-9809-a72191d681bcCited by top-tier papers9
- UTFix: Change Aware Unit Test Repairing using LLMShanto Rahman, Sachit Kuhar, Berk Çirisci, Pranav Garg et al.OOPSLA 2025 · 9 citations
- Retrieval-Augmented Test Generation: How Far Are We?Jiho Shin, Nima Shiri Harzevili, Reem Aleithan, Hadi Hemmati et al.ICSE 2026 · 5 citations
- Generating Project-Specific Test Cases with Requirement Validation IntentionBinhang Qi, Yun Lin, Xinyi Weng, Yuhuan Huang et al.ISSTA 2026 · 3 citations
- REACCEPT: Automated Co-evolution of Production and Test Code Based on Dynamic Validation and Large Language ModelsJianlei Chi, Xiaotian Wang, Yuhan Huang, Lechen Yu et al.ISSTA 2025 · 2 citations
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit TestsJunda Zhao, Shurui Zhou, Eldan CohenISSTA 2026 · 1 citation
Builds on9
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota et al.ICSE 2020 · 96 citations
Related papers
- Test vs Mutant: Adversarial LLM Agents for Robust Unit Test GenerationPengyu Chang, Yixiong Fang, Silin Chen, Yuling Shi et al.ISSTA 2026
- MuMuTestUp: Mutation-Based Multi-agent Test Case UpdateDawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang et al.ISSTA 2026
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
- Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang et al.ISSTA 2026
- Automatic Unit Test Generation for Machine Learning Libraries: How Far Are We?Song Wang, Nishtha Shrestha, Abarna Kucheri Subburaman, Junjie Wang et al.ICSE 2021 · 36 citations
