Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
Qingchuan Ma, Yuhang Wu, Xiawu Zheng, Rongrong Ji
Abstract
In this paper, we aim to establish a simple, effective, and theoretically grounded benchmark for rigorously probing abstract reasoning in Large Language Models (LLMs). To achieve this, we first develop a mathematic framework that defines abstract reasoning as the ability to: (i) extract essential patterns independent of surface representations, and (ii) apply consistent rules to these abstract patterns. Based on this framework, we introduce two novel complementary metrics: Γ measures basic reasoning accuracy, while ∆ quantifies a model's reliance on specific symbols rather than underlying patterns -a key indicator of true abstraction versus mere memorization. To implement this measurement, we design a benchmark: systematic symbol remapping in rule-based tasks, which forces models to demonstrate genuine pattern recognition beyond superficial token matching. Extensive LLM evaluations using this benchmark (commercial API models, 7B-70B, multiagent) reveal:1) critical limitations in non-decimal arithmetic and symbolic reasoning; 2) persistent abstraction gaps despite chain-of-thought prompting; and 3) ∆'s effectiveness in robustly measuring memory dependence by quantifying performance degradation under symbol remapping, particularly highlighting operand-specific memorization. These findings underscore that current LLMs, despite domain-specific strengths, still lack robust abstract reasoning, highlighting key areas for future improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b86170ae-7154-4e31-9f86-9137440fc633Cited by top-tier papers2
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu et al.ICLR 2026 · 20 citations
- A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark GenerationQingchuan Ma, Yuexiao Ma, Yongkang Xie, Tianyu Xie et al.ICML 2026 · 1 citation
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar et al.ICLR 2024 · 114 citations
- A Peek into Token Bias: Large Language Models Are Not Yet Genuine ReasonersBowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang et al.EMNLP 2024 · 27 citations
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio et al.ICLR 2024 · 21 citations
Related papers
- Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer ArithmeticYang Yan, Yu Lu, Renjun Xu, Zhenzhong LanEMNLP 2025 · 1 citation
- Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact GuidanceKai Xiong, Xiao Ding, Ting Liu, Bing Qin et al.NeurIPS 2024
- Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic ComputationZiling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao et al.EMNLP 2025 · 6 citations
- Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping GamesCésar Guerra-Solano, Zhuochun Li, Xiang Lorraine LiEMNLP 2025
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsIman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel et al.ICLR 2025
