Planning a Large Language Model for Static Detection of Runtime Errors in Code Snippets
Smit Patel, Aashish Yadavally, Hridya Dhulipala, Tien N. Nguyen
Abstract
Large Language Models (LLMs) have been excellent in generating and reasoning about source code and natural-language texts. They can recognize patterns, syntax, and semantics in code, making them effective in several software engineering tasks. However, they exhibit weaknesses in reasoning about the program execution. They primarily operate on static code representations, failing to capture the dynamic behavior and state changes that occur during program execution. In this paper, we advance the capabilities of LLMs in reasoning about dynamic program behaviors. We propose Orca, a novel approach that instructs an LLM to autonomously formulate a plan to navigate through a control flow graph (CFG) for predictive execution of (in)complete code snippets. It acts as a predictive interpreter to “execute” the code. In Orca, we guide the LLM to pause at the branching point, focusing on the state of the symbol tables for variables' values, thus minimizing error propagation in the LLM's computation. We instruct the LLM not to stop at each step in its execution plan, resulting the use of only one prompt for the entire predictive interpreter, thus much cost-saving. As a downstream task, we use Orca to statically identify any runtime errors for online code snippets. Early detection of runtime errors and defects in these snippets is crucial to prevent costly fixes later in the development cycle after they were adapted into a codebase. Our empirical evaluation showed that Orca is effective and improves over the state-of-the-art approaches in predicting the execution traces and in static detection of runtime errors.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8cbb9d7b-8408-425c-9947-3c8982090a31Cited by top-tier papers2
- TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language ModelsCuong Chi Le, Cuong Duc Van, Tung Duy Vu, Minh Vu Thai Pham et al.ICSE 2026 · 1 citation
- The Path Not Taken: Duality in Reasoning about Program ExecutionEshgin Hasanov, Md. Mahadi Hassan, Santu Karmaker, Aashish YadavallyACL 2026
Related papers
- Blended Analysis for Predictive ExecutionYi Li, Hridya Dhulipala, Aashish Yadavally, Xiaokai Rong et al.FSE 2025 · 1 citation
- CRISPE: Semantic-Guided Execution Planning and Dynamic Reasoning for Enhancing Code Coverage PredictionHridya Dhulipala, Aashish Yadavally, Smit Soneshbhai Patel, Tien N. NguyenFSE 2025
- TRACED: Execution-aware Pre-training for Source CodeYangruibo Ding, Benjamin Steenhoek, Kexin Pei, Gail E. Kaiser et al.ICSE 2024 · 29 citations
- T-REX: Teaching Large Language Models to Reason with Verbalized Execution SemanticsYan Wang, Ling Ding, Jiechen Sun, Tien N. Nguyen et al.OOPSLA 2026
- Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention InferenceChong Wang, Jianan Liu, Xin Peng, Yang Liu et al.ICSE 2025 · 5 citations
