Retrieval-Augmented Test Generation: How Far Are We?
Jiho Shin, Nima Shiri Harzevili, Reem Aleithan, Hadi Hemmati, Song Wang
Abstract
Retrieval Augmented Generation (RAG) has advanced software engineering tasks but remains underexplored in unit test generation. To bridge this gap, we investigate the efficacy of RAG-based unit test generation for machine learning (ML/DL) APIs and analyze the impact of different knowledge sources on their effectiveness.
We examine three domain-specific sources for RAG: (1) API documentation (official guidelines), (2) GitHub issues (developerreported resolutions), and (3) StackOverflow Q&As (communitydriven solutions). Our study focuses on five widely used Pythonbased ML/DL libraries, TensorFlow, PyTorch, Scikit-learn, Google JAX, and XGBoost, targeting the most-used APIs.
We evaluate four state-of-the-art LLMs: LLMs-GPT-3.5-Turbo, GPT-4o, Mistral MoE 8x22B, and Llama 3.1 405B, across three strategies: basic instruction prompting, Basic RAG, and API-level RAG. Quantitatively, we assess syntactical and dynamic correctness and line coverage. While RAG does not enhance correctness, RAG improves line coverage by 6.5% on average. We found that GitHub issues result in the best improvement in line coverage by providing edge cases from various issues. We also found that these generated unit tests can help detect new bugs. Specifically, 28 bugs were detected, 24 unique bugs were reported to developers, ten were confirmed, four were rejected, and ten are awaiting developers' confirmation.
Our findings highlight RAG's potential in unit test generation for improving test coverage with well-targeted knowledge sources. Future work should focus on retrieval techniques that identify documents with unique program states to optimize RAG-based unit test generation further.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 770d34f3-4e77-4bc9-a8d8-830c0419bd69Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 156 citations
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program RepairWeishi Wang, Yue Wang, Shafiq Joty, Steven C. H. HoiFSE 2023 · 84 citations
Related papers
- LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test GenerationGwihwan Go, Quan Zhang, Chijin Zhou, Zhao Wei et al.ICSE 2026 · 2 citations
- Knowledge-Enhanced Program Repair for Data Science CodeShuyin Ouyang, Jie M. Zhang, Zeyu Sun, Albert Meroño-PeñuelaICSE 2025 · 2 citations
- Are LLMs Correctly Integrated into Software Systems?Yuchen Shao, Yuheng Huang, Jiawei Shen, Lei Ma et al.ICSE 2025 · 4 citations
- Rug: Turbo Llm for Rust Unit Test GenerationXiang Cheng, Fan Sang, Yizhuo Zhai, Xiaokuan Zhang et al.ICSE 2025 · 6 citations
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language ModelsDianshu Liao, Xin Yin, Shidong Pan, Chao Ni et al.ASE 2025 · 2 citations
