Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models
Changshu Liu, Yang Chen, Reyhaneh Jabbarvand
摘要
This paper proposes CES (Code Execution Simulation), a task to evaluate the abilities of LLMs in simulating program execution. Besides measuring the correctness of variable predictions during execution simulation, CES introduces the notion of coherence to determine whether the simulation complies with commonsense execution logic, even if the predicted values along the simulations are incorrect. This enables CES to rule out suspiciously correct output predictions due to reasoning shortcuts, hallucinations, or potential data leakage. CES also introduces a novel metric to measure reasoning consistency across tests with the same or different prime path coverage in a spectrum: strong, weak, and random.
Evaluating 16 LLMs (including three reasoning LLMs) using CES indicates 81.42% coherent execution simulation on HumanEval, 46.92% and 53.08% of which result in correct and incorrect output predictions. Frontier LLMs such as GPT-4 and DeepSeek-R1 have the most incoherent execution reasoning, mostly due to naturallanguage shortcuts. Despite relatively coherent execution simulation, LLMs' reasoning performance across different tests is inconsistent, mostly random (48.87%) or weak (45.37%), potentially explaining their weakness in programming tasks that require path-sensitive program analysis to succeed. We also compare CES with bug prediction/localization/repair, which intuitively requires control-and data-flow awareness. We observe that LLMs rarely incorporate execution reasoning into their analysis for bug-related tasks, and their success is primarily due to inherent pattern-matching capabilities, natural language shortcuts, or data leakage. Without reasoning, there is a threat to the generalizability of LLMs in dealing with unseen bugs or patterns in different contexts. CES can be used to vet the suspicious success of LLMs in these tasks systematically.
• Software and its engineering → Software notations and tools; Correctness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama 等ICML 2024 · 被引用 270 次
相关 Paper
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?Junkai Chen, Zhiyuan Pan, Xing Hu, Zhenhao Li 等ICSE 2025 · 被引用 5 次
- CRISPE: Semantic-Guided Execution Planning and Dynamic Reasoning for Enhancing Code Coverage PredictionHridya Dhulipala, Aashish Yadavally, Smit Soneshbhai Patel, Tien N. NguyenFSE 2025
- Learning to Reason via Program Generation, Emulation, and SearchNathaniel Weir, Muhammad Khalifa, Linlu Qiu, Orion Weller 等NeurIPS 2024 · 被引用 16 次
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng 等ICML 2024 · 被引用 73 次
- Large Language Model Powered Symbolic ExecutionYihe Li, Ruijie Meng, Gregory J. DuckOOPSLA 2025 · 被引用 12 次
