RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
Xinnuo Xu, Rachel Lawrence, Kshitij Dubey, Atharva Pandey, Risa Ueno, Fabian Falck, Aditya V. Nori, Rahul Sharma, Amit Sharma, Javier González
Abstract
Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true "reasoning" or from statistical recall of the training set. Inspired by the ladder of causation (Pearl, 2009) and its three levels (associations, interventions and counterfactuals), this paper introduces RE-IMAGINE: a framework to characterize a hierarchy of reasoning ability in LLMs, alongside a scalable pipeline to generate problem variations across all the levels of the hierarchy. By altering problems in an intermediate symbolic representation, RE-IMAGINE generates arbitrarily many problems that are not solvable using memorization alone. The framework is general and can work across reasoning domains, including math, code, and logic. We demonstrate the type of insights that RE-IMAGINE can generate on four widely-used benchmarks, which we use to evaluate reasoning on several families of LLMs. We observe reductions in performance when the models are queried with problem variations. These assessments indicate a degree of reliance on statistical recall for past performance, and open the door to further research targeting skills across the reasoning hierarchy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e717009-2f52-4687-821f-cd4755c6750fCited by top-tier papers4
- RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning ModelsYunseok Han, Yejoon Lee, Jaeyoung DoICLR 2026 · 10 citations
- Are Language Models Efficient Reasoners? A Perspective from Logic ProgrammingAndreas Opedal, Yanick Zengaffinen, Haruki Shirakami, Clemente Pasti et al.NeurIPS 2025 · 4 citations
- Test of Time: Rethinking Temporal Signal of Benchmark ContaminationTerry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter et al.ACL 2026 · 3 citations
- Omitted Variable Bias in Language Models Under Distribution ShiftVictoria Lin, Louis-Philippe Morency, Eli Ben-MichaelICML 2026 · 1 citation
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- Neuro-Symbolic Data Generation for Math ReasoningZenan Li, Zhi Zhou, Yuan Yao, Xian Zhang et al.NeurIPS 2024 · 35 citations
Related papers
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?Haoang Chi, He Li, Wenjing Yang, Feng Liu et al.NeurIPS 2024 · 124 citations
- Unveiling the Magic of Code Reasoning through Hypothesis Decomposition and AmendmentYuze Zhao, Tianyun Ji, Wenjun Feng, Zhenya Huang et al.ICLR 2025
- NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured NoiseZhi Xu, Yun FuACL 2026
- METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language ModelsPengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen et al.ACL 2026 · 1 citation
