Retrieval-Augmented Test Generation: How Far Are We?
Jiho Shin, Nima Shiri Harzevili, Reem Aleithan, Hadi Hemmati, Song Wang
摘要
Retrieval Augmented Generation (RAG) has advanced software engineering tasks but remains underexplored in unit test generation. To bridge this gap, we investigate the efficacy of RAG-based unit test generation for machine learning (ML/DL) APIs and analyze the impact of different knowledge sources on their effectiveness.
We examine three domain-specific sources for RAG: (1) API documentation (official guidelines), (2) GitHub issues (developerreported resolutions), and (3) StackOverflow Q&As (communitydriven solutions). Our study focuses on five widely used Pythonbased ML/DL libraries, TensorFlow, PyTorch, Scikit-learn, Google JAX, and XGBoost, targeting the most-used APIs.
We evaluate four state-of-the-art LLMs: LLMs-GPT-3.5-Turbo, GPT-4o, Mistral MoE 8x22B, and Llama 3.1 405B, across three strategies: basic instruction prompting, Basic RAG, and API-level RAG. Quantitatively, we assess syntactical and dynamic correctness and line coverage. While RAG does not enhance correctness, RAG improves line coverage by 6.5% on average. We found that GitHub issues result in the best improvement in line coverage by providing edge cases from various issues. We also found that these generated unit tests can help detect new bugs. Specifically, 28 bugs were detected, 24 unique bugs were reported to developers, ten were confirmed, four were rejected, and ten are awaiting developers' confirmation.
Our findings highlight RAG's potential in unit test generation for improving test coverage with well-targeted knowledge sources. Future work should focus on retrieval techniques that identify documents with unique program states to optimize RAG-based unit test generation further.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 被引用 531 次
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 被引用 156 次
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang 等FSE 2024 · 被引用 89 次
- RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program RepairWeishi Wang, Yue Wang, Shafiq Joty, Steven C. H. HoiFSE 2023 · 被引用 84 次
相关 Paper
- LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test GenerationGwihwan Go, Quan Zhang, Chijin Zhou, Zhao Wei 等ICSE 2026 · 被引用 2 次
- Knowledge-Enhanced Program Repair for Data Science CodeShuyin Ouyang, Jie M. Zhang, Zeyu Sun, Albert Meroño-PeñuelaICSE 2025 · 被引用 2 次
- Are LLMs Correctly Integrated into Software Systems?Yuchen Shao, Yuheng Huang, Jiawei Shen, Lei Ma 等ICSE 2025 · 被引用 4 次
- Rug: Turbo Llm for Rust Unit Test GenerationXiang Cheng, Fan Sang, Yizhuo Zhai, Xiaokuan Zhang 等ICSE 2025 · 被引用 6 次
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language ModelsDianshu Liao, Xin Yin, Shidong Pan, Chao Ni 等ASE 2025 · 被引用 2 次
