Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)
Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang, Xiao Chu, Jianyi Zhou, Guangtai Liang, Qianxiang Wang, Dong Wang
摘要
Automated unit test generation has recently benefited from advances in large language models (LLMs), yet our industrial deployments reveal a persistent gap between promising research results and practical usability. In real-world projects with complex frameworks and cross-file dependencies, LLM-generated tests frequently fail to compile, require costly manual repair, or provide unstable coverage improvements. This paper reports our experience in designing, deploying, and evaluating CATGen , a context-aware workflow for LLM-based unit test generation, informed by repeated industrial failures and refinements. Rather than relying on LLMs to infer incomplete project context, we found that compilation robustness critically depends on making project-level dependencies explicit, stabilizing test class scaffolding, and replacing iterative LLM-based repair with lightweight static analysis. These experience-driven insights shaped CATGen’s multi-stage design, which combines structured context retrieval, deterministic test skeleton construction, and program analysis–based post-processing. We evaluate CATGen on real-world complex focal methods from proprietary industrial projects and additionally on the Defects4J benchmark to assess generalizability. Across both settings, CATGen substantially improves compilation success and structural coverage while significantly reducing generation time and token consumption compared to existing LLM-based approaches. Our results demonstrate that reliable LLM-based unit test generation in practice depends less on prompt engineering alone and more on systematic engineering support grounded in real-world development constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang 等FSE 2024 · 被引用 89 次
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney 等ICSE 2023 · 被引用 45 次
相关 Paper
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language ModelsDianshu Liao, Xin Yin, Shidong Pan, Chao Ni 等ASE 2025 · 被引用 2 次
- STRUT: Structured Seed Case Guided Unit Test Generation for C Programs using LLMsJinwei Liu, Chao Li, Rui Chen, Shaofeng Li 等ISSTA 2025 · 被引用 6 次
- Enhancing LLM-Based Bug Reproduction via Code Entity Retrieval and Test Case RepairHao Ding, Yanjie Jiang, Yuxia Zhang, Hui LiuISSTA 2026
- Knowledge Matters: Injecting Project and Testing Knowledge into LLM-based Unit Test GenerationAnji Li, Mingwei Liu, Zhenxi Chen, Zheng Pei 等ICSE 2026 · 被引用 2 次
- FlakyGuard: Automatically Fixing Flaky Tests at Industry ScaleChengpeng Li, Farnaz Behrang, August Shi, Peng LiuASE 2025 · 被引用 1 次
