CRISPE: Semantic-Guided Execution Planning and Dynamic Reasoning for Enhancing Code Coverage Prediction
Hridya Dhulipala, Aashish Yadavally, Smit Soneshbhai Patel, Tien N. Nguyen
Abstract
While LLMs excel in understanding source code and descriptive texts for tasks like code generation, code completion, etc., they exhibit weaknesses in predicting dynamic program behavior, such as code coverage and runtime error detection, which typically require program execution. Aiming to advance the capability of LLMs in reasoning and predicting the program behavior at runtime, we present CRISPE (short for Coverage Rationalization and Intelligent Selection ProcedurE), a novel approach for code coverage prediction. CRISPE guides an LLM in simulating program execution via an execution plan based on two key factors: (1) program semantics of each statement type, and (2) the observation of the set of covered statements at the current “execution” step relative to all feasible code coverage options. We formulate code coverage prediction as a process of semantic-guided execution-based planning, where feasible coverage options are utilized to assess whether the LLM is heading in the correct reasoning. We enhance the traditional generative task with the retrieval-based framework on feasible options of code coverage. Our experimental results show that CRISPE achieves high accuracy in coverage prediction in terms of both exact-match and statement-match coverage metrics, improving over the baselines. We also show that with semantic-guiding and dynamic reasoning from CRISPE, the LLM generates more correct planning steps. To demonstrate CRISPE’s usefulness, we used it in the downstream task of statically detecting runtime error(s) in incomplete code snippets with the given inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a14ed7ef-dcbc-414a-8360-78eb0f7f2e2eCited by top-tier papers1
Ask how each one uses itBuilds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Boosting coverage-based fault localization via graph-based representation learningYiling Lou, Qihao Zhu, Jinhao Dong, Xia Li et al.FSE 2021 · 157 citations
- Full-Speed Fuzzing: Reducing Fuzzing Overhead through Coverage-Guided TracingStefan Nagy, Matthew HicksS&P 2019 · 156 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng et al.ICML 2024 · 73 citations
Related papers
- Blended Analysis for Predictive ExecutionYi Li, Hridya Dhulipala, Aashish Yadavally, Xiaokai Rong et al.FSE 2025 · 1 citation
- Planning a Large Language Model for Static Detection of Runtime Errors in Code SnippetsSmit Patel, Aashish Yadavally, Hridya Dhulipala, Tien N. NguyenICSE 2025 · 1 citation
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language ModelsChangshu Liu, Yang Chen, Reyhaneh JabbarvandICSE 2026
- TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language ModelsCuong Chi Le, Cuong Duc Van, Tung Duy Vu, Minh Vu Thai Pham et al.ICSE 2026 · 1 citation
- T-REX: Teaching Large Language Models to Reason with Verbalized Execution SemanticsYan Wang, Ling Ding, Jiechen Sun, Tien N. Nguyen et al.OOPSLA 2026
