CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, Bernhard Schölkopf
Abstract
The ability to perform causal reasoning is widely considered a core feature of intelligence. In this work, we investigate whether large language models (LLMs) can coherently reason about causality. Much of the existing work in natural language processing (NLP) focuses on evaluating commonsense causal reasoning in LLMs, thus failing to assess whether a model can perform causal inference in accordance with a set of well-defined formal rules. To address this, we propose a new NLP task, causal inference in natural language, inspired by the "causal inference engine" postulated by Judea Pearl et al. We compose a large dataset, CLADDER, with 10K samples: based on a collection of causal graphs and queries (associational, interventional, and counterfactual), we obtain symbolic questions and ground-truth answers, through an oracle causal inference engine. These are then translated into natural language. We evaluate multiple LLMs on our dataset, and we introduce and evaluate a bespoke chain-of-thought prompting strategy, CAUSALCOT. We show that our task is highly challenging for LLMs, and we conduct an in-depth analysis to gain deeper insights into the causal reasoning abilities of LLMs. 1 * Main contributors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1e3b7e7-48e1-45d8-8d3c-9e3cdfc6c853Cited by top-tier papers17
- Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao et al.ICML 2024 · 27 citations
- Discovery of the Hidden World with Large Language ModelsChenxi Liu, Yongqiang Chen, Tongliang Liu, Mingming Gong et al.NeurIPS 2024 · 26 citations
- FoNE: Precise Single-Token Number Embeddings via Fourier FeaturesTianyi Zhou, Deqing Fu, Mahdi Soltanolkotabi, Robin Jia et al.ICLR 2026 · 24 citations
- Limits of Transformer Language Models on Learning to Compose AlgorithmsJonathan Thomm, Giacomo Camposampiero, Aleksandar Terzic, Michael Hersche et al.NeurIPS 2024 · 16 citations
- Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News DetectionChaowei Zhang, Zongling Feng, Zewei Zhang, Jipeng Qiang et al.AAAI 2025 · 13 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
Related papers
- NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured NoiseZhi Xu, Yun FuACL 2026
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- Mitigating Hallucinations in Large Language Models via Causal ReasoningYuangang Li, Yiqing Shen, Yi Nian, Jiechao Gao et al.AAAI 2026 · 1 citation
- CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language ModelsYuefei Chen, Vivek K. Singh, Jing Ma, Ruixiang TangAAAI 2026 · 1 citation
- Reasoning over Uncertain Text by Generative Large Language ModelsAliakbar Nafar, Kristen Brent Venable, Parisa KordjamshidiAAAI 2025 · 13 citations
