LLMs Lean on Priors, Not Programming Language Semantics
Aditya Thimmaiah, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric
摘要
Recent work asks whether large language models (LLMs) condition their reasoning on explicit rules rather than statistical regularities from pretraining. Program execution provides a canonical instance: formal semantics define behavior through symbolic transition rules that can be systematically altered under distribution shift. We investigate whether LLMs can condition their reasoning on formal semantics through program execution and introduce PLSEMANTICSBENCH, pairing featherweight C programs with two semantic systems-small-step operational semantics and K semantics-and probing four capabilities: composing rules for final states, selecting rules when state is unmutated, sustaining such conditioning over long traces, and following supplied rules under novel semantics. To decouple semantic reasoning from syntactic familiarity, we redefine familiar operators to induce symbol-meaning conflict and introduce novel symbols defined only through the supplied rules, and stress-test models on Human-Written, LLM-Translated, and Fuzzer-Generated splits with increasing structural complexity. Across 11 frontier LLMs, strong finalstate accuracy under standard semantics (up to 90%) drops sharply-by as much as 40-60% points-under semantic mutations and increasing structural complexity. Only a handful of models achieve non-zero long-horizon conditioning accuracy, and even the best systems reach just 35%. Together, these results suggest that contemporary LLMs often rely on pretrained lexical associations rather than systematically conditioning on supplied formal rules. PLSEMANTICSBENCH is publicly available at https://EngineeringS oftware.github.io/PLSemanticsBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama 等ICML 2024 · 被引用 270 次
- CodeAlchemist: Semantics-Aware Code Generation to Find Vulnerabilities in JavaScript EnginesHyungSeok Han, DongHyeon Oh, Sang Kil ChaNDSS 2019 · 被引用 178 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
- Measuring the Impact of Programming Language DistributionGabriel Orlanski, Kefan Xiao, Xavier Garcia, Jeffrey Hui 等ICML 2023 · 被引用 49 次
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 被引用 34 次
相关 Paper
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence CheckingAnjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen 等EMNLP 2025
- The Path Not Taken: Duality in Reasoning about Program ExecutionEshgin Hasanov, Md. Mahadi Hassan, Santu Karmaker, Aashish YadavallyACL 2026
- Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code UnderstandingAdam Storek, Mukur Gupta, Samira Hajizadeh, Prashast Srivastava 等ACL 2026 · 被引用 4 次
- MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation DatasetWeiqi Wang, Yangqiu SongACL 2025
