The Path Not Taken: Duality in Reasoning about Program Execution
Eshgin Hasanov, Md. Mahadi Hassan, Santu Karmaker, Aashish Yadavally
Abstract
Large language models (LLMs) have shown remarkable capabilities across diverse coding tasks. However, their adoption requires a true understanding of program execution rather than relying on surface-level patterns. Existing benchmarks primarily focus on predicting program properties tied to specific inputs (e.g., code coverage, program outputs). As a result, they provide a narrow view of dynamic code reasoning and are prone to data contamination. We argue that understanding program execution requires evaluating its inherent duality through two complementary reasoning tasks: (i) predicting a program's observed behavior for a given input, and (ii) inferring how the input must be mutated toward a specific behavioral objective. Both tasks jointly probe a model's causal understanding of execution flow. We instantiate this duality in DexBench, a benchmark comprising 445 paired instances, and evaluate 13 LLMs. Our results demonstrate that dual-path reasoning provides a robust and discriminative proxy for dynamic code understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ffeae10-16b2-4f47-83ac-42e943ce4876Builds on14
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 836 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 45 citations
Related papers
- DRPBench: Evaluating LLMs in Concurrent Code Comprehension via Fine-grained Data Race PredictionYuqi Guo, Siwei Wei, Yan CaiICML 2026
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?Junkai Chen, Zhiyuan Pan, Xing Hu, Zhenhao Li et al.ICSE 2025 · 5 citations
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence CheckingAnjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen et al.EMNLP 2025
- Blended Analysis for Predictive ExecutionYi Li, Hridya Dhulipala, Aashish Yadavally, Xiaokai Rong et al.FSE 2025 · 1 citation
