TestTailor: Generating High-Coverage Tests via Path-Proximal Tests with LLMs
Xiaoxuan Zhou, Yiling Lou, Jinhao Dong, Dan Hao
Abstract
Automated unit testing is essential for ensuring software quality. Achieving high code coverage through automated unit test generation remains challenging, especially for hard-to-cover branches guarded by complex or deeply nested conditions. Traditional search-based approaches often stagnate at fitness plateaus, while recent LLM-based techniques provide mostly coarse-grained prompts, leaving models to guess how to reach uncovered targets. To address these limitations, we present TestTailor, a neuro-symbolic framework that exploits fine-grained, path-oriented guidance to guide LLM-based test generation. The key idea is to exploit path-proximal tests (i.e., existing test cases whose execution paths closely resemble the target uncovered path) and to analyze their divergence points. By combining this analysis with symbolic constraints (i.e., constraints collected from the target uncovered path using symbolic execution), TestTailor derives actionable path guidance and encodes them into concise prompts that tell the LLM not only what remains uncovered, but also how to reach it. We evaluate TestTailor on the widely used CODAMOSA benchmark comprising 486 Python modules. Results show that TestTailor consistently outperforms state-of-the-art baselines, improving statement coverage by 5.01% and branch coverage by 4.17% on average compared to the best baseline CoverUp, while incurring only about 40% of CoverUp's API cost. Against the hybrid LLM-search-based technique CODAMOSA, TestTailor achieves even larger gains of 12.23% and 12.54% in statement and branch coverage, respectively. Moreover, TestTailor attains the highest coverage accuracy (85.2% vs. 75.3% for CoverUp and 63.8% for TELPA), and demonstrates robustness across different LLM backbones. These results highlight that TestTailor transforms vague coverage goals into precise path-level instructions, enabling LLMs to generate high-coverage test suites more efficiently and accurately.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2d92493b-2de0-476a-a53a-ee95b80badf8Related papers
- CoverUp: Effective High Coverage Test Generation for PythonJuan Altmayer Pizzorno, Emery D. BergerFSE 2025 · 13 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- Towards Understanding the Effectiveness of Large Language Models on Directed Test Input GenerationZongze Jiang, Ming Wen, Jialun Cao, Xuanhua Shi et al.ASE 2024 · 8 citations
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang et al.FSE 2024 · 68 citations
- ConUT: Condition-Aware Test Generation for Complex Java CodeRuiguo Yu, Ruiqi Dong, Xi Xiao, Xiaogang Zhu et al.ISSTA 2026
