LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test Generation
Gwihwan Go, Quan Zhang, Chijin Zhou, Zhao Wei, Yu Jiang
Abstract
Automated unit test generation is essential for robust software development, yet existing approaches struggle to generalize across multiple programming languages and operate within real-time development. While Large Language Models (LLMs) offer a promising solution, their ability to generate high coverage test code depends on prompting a concise context of the focal method. Current solutions, such as Retrieval-Augmented Generation, either rely on imprecise similarity-based searches or demand the creation of costly, language-specific static analysis pipelines. To address this gap, we present LspRag, a framework for concise-context retrieval tailored for real-time, language-agnostic unit test generation. LspRag leverages off-the-shelf Language Server Protocol (LSP) back-ends to supply LLMs with precise symbol definitions and references in real time. By reusing mature LSP servers, LspRag provides an LLM with language-aware context retrieval, requiring minimal per-language engineering effort. We evaluated LspRag on open-source projects spanning Java, Go, and Python. Compared to the best performance of baselines, LspRag increased line coverage by up to 174.55% for Golang, 213.31% for Java, and 31.57% for Python.
• Software and its engineering → Software testing and debugging; • Computing methodologies → Neural networks; Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db83af10-39a5-4c0b-b435-b4a0cba0cd79Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo et al.ACL 2022 · 208 citations
Related papers
- Retrieval-Augmented Test Generation: How Far Are We?Jiho Shin, Nima Shiri Harzevili, Reem Aleithan, Hadi Hemmati et al.ICSE 2026 · 5 citations
- Enhancing LLM's Ability to Generate More Repository-Aware Unit Tests Through Precise Context InjectionXin Yin, Chao Ni, Xinrui Li, Liushan Chen et al.ASE 2025 · 2 citations
- STRUT: Structured Seed Case Guided Unit Test Generation for C Programs using LLMsJinwei Liu, Chao Li, Rui Chen, Shaofeng Li et al.ISSTA 2025 · 6 citations
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language ModelsDianshu Liao, Xin Yin, Shidong Pan, Chao Ni et al.ASE 2025 · 2 citations
- Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang et al.ISSTA 2026
