EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
Anjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen, Yuhui Zhang, Ziheng Wang, Yuan Liu, Thiago S. F. X. Teixeira, Diyi Yang, Ke Wang, Alex Aiken
Abstract
As large language models (LLMs) become integral to code-related tasks, a central question emerges: Do LLMs truly understand program semantics? We introduce EquiBench, a new benchmark for evaluating LLMs through equivalence checking, i.e., determining whether two programs produce identical outputs for all possible inputs. Unlike prior code generation benchmarks, this task directly tests a model's ability to reason about program semantics. EquiBench consists of 2400 program pairs across four languages and six categories. These pairs are generated through program analysis, compiler scheduling, and superoptimization, ensuring high-confidence labels, nontrivial difficulty, and full automation. We evaluate 19 state-of-the-art LLMs and find that in the most challenging categories, the best accuracies are 63.8% and 76.2%, only modestly above the 50% random baseline. Further analysis reveals that models often rely on syntactic similarity rather than exhibiting robust reasoning about program semantics, highlighting current limitations. Our code and dataset are publicly available at https://github.com/Anjiang-Wei/ equibench
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41e742a1-f65e-4a34-a062-8a3f74f1027fCited by top-tier papers6
- MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and RepairAli Reza Ibrahimzada, Brandon Paulsen, Reyhaneh Jabbarvand, Joey Dodds et al.ICML 2026 · 9 citations
- SciCoQA: Quality Assurance for Scientific Paper-Code AlignmentTim Baumgärtner, Iryna GurevychACL 2026 · 5 citations
- InnoGym: Benchmarking the Innovation Potential of AI AgentsJintian Zhang, Kewei Xu, Jingsheng Zheng, Zhuoyun Yu et al.ICLR 2026 · 4 citations
- LLMs Lean on Priors, Not Programming Language SemanticsAditya Thimmaiah, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li et al.ICML 2026 · 2 citations
- Don't Force the Fit: Bounded Log-Likelihood Loss for Enhanced Reasoning in Large Language ModelsFeng Zhao, Hong Zhang, Yu Yang, Ruilin Zhao et al.ICML 2026
Builds on16
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- Learning Performance-Improving Code EditsAlexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon et al.ICLR 2024 · 141 citations
- Large Language Models are Visual Reasoning CoordinatorsLiangyu Chen, Bo Li, Sheng Shen, Jingkang Yang et al.NeurIPS 2023 · 108 citations
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 102 citations
Related papers
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 5 citations
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodeLingfei Zeng, Fengdi Che, Xuhan Huang, Fei Ye et al.ICLR 2026 · 8 citations
- GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMsShixian Luo, Zhu zezhou, Yu Yuan, Yuncheng Yang et al.ICLR 2026 · 15 citations
- ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API CallsRenbiao Liu, Chao-Zeng Ma, Anqi Li, Hui Sun et al.AAAI 2026
