Towards Understanding the Effectiveness of Large Language Models on Directed Test Input Generation
Zongze Jiang, Ming Wen, Jialun Cao, Xuanhua Shi, Hai Jin
Abstract
Automatic testing has garnered significant attention and success over the past few decades. Techniques such as unit testing and coverage-guided fuzzing have revealed numerous critical software bugs and vulnerabilities. However, a long-standing, formidable challenge for existing techniques is how to achieve higher testing coverage. Constraint-based techniques, such as symbolic execution and concolic testing, have been well-explored and integrated into the existing approaches. With the popularity of Large Language Models (LLMs), recent research efforts to design tailored prompts to generate inputs that can reach more uncovered target branches. However, the effectiveness of using LLMs for generating such directed inputs and the comparison with the proven constraint-based solutions has not been systematically explored.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 64b42ae7-180c-4a99-9e1c-64e8d09e6adaCited by top-tier papers11
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language ModelsDianshu Liao, Xin Yin, Shidong Pan, Chao Ni et al.ASE 2025 · 2 citations
- Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMsYifan Xia, Zichen Xie, Peiyu Liu, Kangjie Lu et al.ISSTA 2025 · 2 citations
- Comprehend, Imitate, and then Update: Unleashing the Power of LLMs in Test Suite EvolutionTangzhi Xu, Jianhan Liu, Yuan Yao, Cong Li et al.ASE 2025 · 1 citation
- Evaluating LLM-Based Regression Test GenerationJing Liu, Seongmin Lee, Eleonora Losiouk, Marcel BöhmeFSE 2026 · 1 citation
- WEDGE: Synthesizing Performance Constraints for Evaluating and Improving Code EfficiencyJun Yang, Cheng-Chi Wang, Bogdan Alexandru Stoica, Kexin PeiNeurIPS 2025
Related papers
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- Agentic Concolic ExecutionZhengxiong Luo, Huan Zhao, Dylan Wolff, Cristian Cadar et al.S&P 2026 · 17 citations
- Cottontail: Large Language Model-Driven Concolic Execution for Highly Structured Test Input GenerationHaoxin Tu, Seongmin Lee, Yuxian Li, Peng Chen et al.S&P 2026 · 22 citations
- PALM: Synergizing Program Analysis and LLMs to Enhance Rust Unit Test CoverageBei Chu, Yang Feng, Kui Liu, Hange Shi et al.ASE 2025 · 3 citations
- TestTailor: Generating High-Coverage Tests via Path-Proximal Tests with LLMsXiaoxuan Zhou, Yiling Lou, Jinhao Dong, Dan HaoFSE 2026
