Evaluating Relational Reasoning in LLMs with REL
Lukas Fesser, Yasha Ektefaie, Ada Fang, Sham Kakade, Marinka Zitnik
Abstract
Relational reasoning is the ability to infer relations that jointly bind multiple entities, attributes, or variables. While this capability is essential for scientific reasoning, most existing evaluations of relational reasoning in large language models focus on structured inputs such as tables, graphs, or synthetic relational tasks, and do not isolate the sources of difficulty that arise from higher-arity relational binding. We study this problem through the lens of Relational Complexity (RC) , defined as the minimum number of independent entities or operands that must be simultaneously bound to apply a relation. RC provides a principled way to vary reasoning difficulty independently of confounders such as input size, vocabulary, and representational choices. Building on RC, we introduce REL, a generative benchmark framework spanning algebra, chemistry, and biology that varies RC within each domain. Evaluating frontier LLMs, we observe a consistent and monotonic degradation in performance as RC increases, even when the total number of entities is held fixed. This failure mode persists under increased test-time compute and with in-context learning, suggesting a limitation tied to the arity of the required relational binding rather than insufficient inference steps or exposure to examples. Our results identify a well-defined regime of higher-arity reasoning in which current models struggle and motivate revisiting reasoning benchmarks through the lens of relational complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c292db54-64ee-4aff-b8f5-3f2fd852f933Builds on13
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 499 citations
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question AnsweringXiaoxin He, Yijun Tian, Yifei Sun, Nitesh V. Chawla et al.NeurIPS 2024 · 384 citations
- Talk like a Graph: Encoding Graphs for Large Language ModelsBahare Fatemi, Jonathan Halcrow, Bryan PerozziICLR 2024 · 194 citations
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li et al.ICML 2024 · 184 citations
Related papers
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsMohammad Ramezanali, Mo Vazifeh, Paolo SantiEMNLP 2025
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl et al.ICLR 2026 · 5 citations
- Can LLMs Reason Structurally? Benchmarking via the lens of Data StructuresYu He, Yingxi Li, Colin White, Ellen VitercikICML 2026 · 3 citations
- ReCogLab: a framework testing relational reasoning & cognitive hypotheses on LLMsAndrew Liu, Henry Prior, Gargi Balasubramaniam, Rivka Moroshko et al.ICLR 2025
