LLMs Lean on Priors, Not Programming Language Semantics
Aditya Thimmaiah, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric
Abstract
Recent work asks whether large language models (LLMs) condition their reasoning on explicit rules rather than statistical regularities from pretraining. Program execution provides a canonical instance: formal semantics define behavior through symbolic transition rules that can be systematically altered under distribution shift. We investigate whether LLMs can condition their reasoning on formal semantics through program execution and introduce PLSEMANTICSBENCH, pairing featherweight C programs with two semantic systems-small-step operational semantics and K semantics-and probing four capabilities: composing rules for final states, selecting rules when state is unmutated, sustaining such conditioning over long traces, and following supplied rules under novel semantics. To decouple semantic reasoning from syntactic familiarity, we redefine familiar operators to induce symbol-meaning conflict and introduce novel symbols defined only through the supplied rules, and stress-test models on Human-Written, LLM-Translated, and Fuzzer-Generated splits with increasing structural complexity. Across 11 frontier LLMs, strong finalstate accuracy under standard semantics (up to 90%) drops sharply-by as much as 40-60% points-under semantic mutations and increasing structural complexity. Only a handful of models achieve non-zero long-horizon conditioning accuracy, and even the best systems reach just 35%. Together, these results suggest that contemporary LLMs often rely on pretrained lexical associations rather than systematically conditioning on supplied formal rules. PLSEMANTICSBENCH is publicly available at https://EngineeringS oftware.github.io/PLSemanticsBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4747935-c9e4-477a-96b7-593a32778f0dBuilds on12
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- CodeAlchemist: Semantics-Aware Code Generation to Find Vulnerabilities in JavaScript EnginesHyungSeok Han, DongHyeon Oh, Sang Kil ChaNDSS 2019 · 178 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- Measuring the Impact of Programming Language DistributionGabriel Orlanski, Kefan Xiao, Xavier Garcia, Jeffrey Hui et al.ICML 2023 · 49 citations
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 34 citations
Related papers
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence CheckingAnjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen et al.EMNLP 2025
- The Path Not Taken: Duality in Reasoning about Program ExecutionEshgin Hasanov, Md. Mahadi Hassan, Santu Karmaker, Aashish YadavallyACL 2026
- Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code UnderstandingAdam Storek, Mukur Gupta, Samira Hajizadeh, Prashast Srivastava et al.ACL 2026 · 4 citations
- MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation DatasetWeiqi Wang, Yangqiu SongACL 2025
