Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)
Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang, Xiao Chu, Jianyi Zhou, Guangtai Liang, Qianxiang Wang, Dong Wang
Abstract
Automated unit test generation has recently benefited from advances in large language models (LLMs), yet our industrial deployments reveal a persistent gap between promising research results and practical usability. In real-world projects with complex frameworks and cross-file dependencies, LLM-generated tests frequently fail to compile, require costly manual repair, or provide unstable coverage improvements. This paper reports our experience in designing, deploying, and evaluating CATGen , a context-aware workflow for LLM-based unit test generation, informed by repeated industrial failures and refinements. Rather than relying on LLMs to infer incomplete project context, we found that compilation robustness critically depends on making project-level dependencies explicit, stabilizing test class scaffolding, and replacing iterative LLM-based repair with lightweight static analysis. These experience-driven insights shaped CATGen’s multi-stage design, which combines structured context retrieval, deterministic test skeleton construction, and program analysis–based post-processing. We evaluate CATGen on real-world complex focal methods from proprietary industrial projects and additionally on the Defects4J benchmark to assess generalizability. Across both settings, CATGen substantially improves compilation success and structural coverage while significantly reducing generation time and token consumption compared to existing LLM-based approaches. Our results demonstrate that reliable LLM-based unit test generation in practice depends less on prompt engineering alone and more on systematic engineering support grounded in real-world development constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c024c9bb-be9b-4127-9394-ce779fcbb797Builds on14
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota et al.ICSE 2020 · 96 citations
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 92 citations
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney et al.ICSE 2023 · 45 citations
Related papers
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language ModelsDianshu Liao, Xin Yin, Shidong Pan, Chao Ni et al.ASE 2025 · 2 citations
- STRUT: Structured Seed Case Guided Unit Test Generation for C Programs using LLMsJinwei Liu, Chao Li, Rui Chen, Shaofeng Li et al.ISSTA 2025 · 6 citations
- Enhancing LLM-Based Bug Reproduction via Code Entity Retrieval and Test Case RepairHao Ding, Yanjie Jiang, Yuxia Zhang, Hui LiuISSTA 2026
- Knowledge Matters: Injecting Project and Testing Knowledge into LLM-based Unit Test GenerationAnji Li, Mingwei Liu, Zhenxi Chen, Zheng Pei et al.ICSE 2026 · 2 citations
- FlakyGuard: Automatically Fixing Flaky Tests at Industry ScaleChengpeng Li, Farnaz Behrang, August Shi, Peng LiuASE 2025 · 1 citation
