Blended Analysis for Predictive Execution
Yi Li, Hridya Dhulipala, Aashish Yadavally, Xiaokai Rong, Shaohua Wang, Tien N. Nguyen
Abstract
Although Large Language Models (LLMs) are highly proficient in understanding source code and descriptive texts, they have limitations in reasoning on dynamic program behaviors, such as execution trace and code coverage prediction, and runtime error prediction, which usually require actual program execution. To advance the ability of LLMs in predicting dynamic behaviors, we leverage the strengths of both approaches, Program Analysis (PA) and LLM, in building PredEx, a predictive executor for Python. Our principle is a blended analysis between PA and LLM to use PA to guide the LLM in predicting execution traces. We break down the task of predictive execution into smaller sub-tasks and leverage the deterministic nature when an execution order can be deterministically decided. When it is not certain, we use predictive backward slicing per variable, i.e., slicing the prior trace to only the parts that affect each variable separately breaks up the valuation prediction into significantly simpler problems. Our empirical evaluation on real-world datasets shows that PredEx achieves 31.5–47.1% relatively higher accuracy in predicting full execution traces than the state-of-the-art models. It also produces 8.6–53.7% more correct execution trace prefixes than those baselines. In predicting next executed statements, its relative improvement over the baselines is 15.7–102.3%. Finally, we show PredEx’s usefulness in two tasks: static code coverage analysis and static prediction of run-time errors for (in)complete code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Stack Overflow Considered Harmful? The Impact of Copy&Paste on Android Application SecurityFelix Fischer, Konstantin Böttinger, Huang Xiao, Christian Stransky et al.S&P 2017 · 293 citations
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng et al.ICML 2024 · 73 citations
- TRACED: Execution-aware Pre-training for Source CodeYangruibo Ding, Benjamin Steenhoek, Kexin Pei, Gail E. Kaiser et al.ICSE 2024 · 29 citations
Related papers
- T-REX: Teaching Large Language Models to Reason with Verbalized Execution SemanticsYan Wang, Ling Ding, Jiechen Sun, Tien N. Nguyen et al.OOPSLA 2026
- Treefix: Enabling Execution with a Tree of PrefixesBeatriz Souza, Michael PradelICSE 2025
- CRISPE: Semantic-Guided Execution Planning and Dynamic Reasoning for Enhancing Code Coverage PredictionHridya Dhulipala, Aashish Yadavally, Smit Soneshbhai Patel, Tien N. NguyenFSE 2025
- Planning a Large Language Model for Static Detection of Runtime Errors in Code SnippetsSmit Patel, Aashish Yadavally, Hridya Dhulipala, Tien N. NguyenICSE 2025 · 1 citation
- The Path Not Taken: Duality in Reasoning about Program ExecutionEshgin Hasanov, Md. Mahadi Hassan, Santu Karmaker, Aashish YadavallyACL 2026
