CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, Bernhard Schölkopf
摘要
The ability to perform causal reasoning is widely considered a core feature of intelligence. In this work, we investigate whether large language models (LLMs) can coherently reason about causality. Much of the existing work in natural language processing (NLP) focuses on evaluating commonsense causal reasoning in LLMs, thus failing to assess whether a model can perform causal inference in accordance with a set of well-defined formal rules. To address this, we propose a new NLP task, causal inference in natural language, inspired by the "causal inference engine" postulated by Judea Pearl et al. We compose a large dataset, CLADDER, with 10K samples: based on a collection of causal graphs and queries (associational, interventional, and counterfactual), we obtain symbolic questions and ground-truth answers, through an oracle causal inference engine. These are then translated into natural language. We evaluate multiple LLMs on our dataset, and we introduce and evaluate a bespoke chain-of-thought prompting strategy, CAUSALCOT. We show that our task is highly challenging for LLMs, and we conduct an in-depth analysis to gain deeper insights into the causal reasoning abilities of LLMs. 1 * Main contributors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao 等ICML 2024 · 被引用 27 次
- Discovery of the Hidden World with Large Language ModelsChenxi Liu, Yongqiang Chen, Tongliang Liu, Mingming Gong 等NeurIPS 2024 · 被引用 26 次
- FoNE: Precise Single-Token Number Embeddings via Fourier FeaturesTianyi Zhou, Deqing Fu, Mahdi Soltanolkotabi, Robin Jia 等ICLR 2026 · 被引用 24 次
- Limits of Transformer Language Models on Learning to Compose AlgorithmsJonathan Thomm, Giacomo Camposampiero, Aleksandar Terzic, Michael Hersche 等NeurIPS 2024 · 被引用 16 次
- Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News DetectionChaowei Zhang, Zongling Feng, Zewei Zhang, Jipeng Qiang 等AAAI 2025 · 被引用 13 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi 等ICLR 2020 · 被引用 521 次
相关 Paper
- NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured NoiseZhi Xu, Yun FuACL 2026
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff 等ICLR 2024 · 被引用 186 次
- Mitigating Hallucinations in Large Language Models via Causal ReasoningYuangang Li, Yiqing Shen, Yi Nian, Jiechao Gao 等AAAI 2026 · 被引用 1 次
- CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language ModelsYuefei Chen, Vivek K. Singh, Jing Ma, Ruixiang TangAAAI 2026 · 被引用 1 次
- Reasoning over Uncertain Text by Generative Large Language ModelsAliakbar Nafar, Kristen Brent Venable, Parisa KordjamshidiAAAI 2025 · 被引用 13 次
