LLMs Are Prone to Fallacies in Causal Inference
Nitish Joshi, Abulhair Saparov, Yixin Wang, He He
Abstract
Recent work shows that causal facts can be effectively extracted from LLMs through prompting, facilitating the creation of causal graphs for causal inference tasks. However, it is unclear if this success is limited to explicitly-mentioned causal facts in the pretraining data which the model can memorize. Thus, this work investigates: Can LLMs infer causal relations from other relational data in text? To disentangle the role of memorized causal facts vs inferred causal relations, we finetune LLMs on synthetic data containing temporal, spatial and counterfactual relations, and measure whether the LLM can then infer causal relations. We find that: (a) LLMs are susceptible to inferring causal relations from the order of two entity mentions in text (e.g. X mentioned before Y implies X causes Y); (b) if the order is randomized, LLMs still suffer from the post hoc fallacy, i.e. X occurs before Y (temporal relation) implies X causes Y. We also find that while LLMs can correctly deduce the absence of causal relations from temporal and spatial relations, they have difficulty inferring causal relations from counterfactuals, questioning their understanding of causality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80ebbc39-7424-4c15-9e98-3e77fd88a928Cited by top-tier papers4
- Causal Discovery and Inference through Next-Token PredictionEivinas Butkus, Nikolaus KriegeskorteNeurIPS 2025 · 3 citations
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMsKarin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu et al.ACL 2025
- Spiked-CFR: Causal Representation Learning from LLMs via Wasserstein Projection PursuitFan Wang, Hengyu Yue, Yu Bowen, Weiming Liu et al.ICML 2026
- LLMs Struggle to Balance Reasoning and World Knowledge in Causal Narrative UnderstandingKhurram Yamin, Shantanu Gupta, Gaurav R. Ghosal, Zachary C. Lipton et al.ICLR 2026
Builds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Evaluating the Moral Beliefs Encoded in LLMsNino Scherrer, Claudia Shi, Amir Feder, David M. BleiNeurIPS 2023 · 316 citations
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- Teaching Arithmetic to Small TransformersNayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee et al.ICLR 2024 · 128 citations
Related papers
- On the Reliability of Large Language Models for Causal DiscoveryTao Feng, Lizhen Qu, Niket Tandon, Zhuang Li et al.ACL 2025
- Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMsYihua Zhu, Qianying Liu, Jiaxin Wang, Fei Cheng et al.ACL 2026 · 1 citation
- Causal Order: The Key to Leveraging Imperfect Experts in Causal InferenceAniket Vashishtha, Abbavaram Gowtham Reddy, Abhinav Kumar, Saketh Bachu et al.ICLR 2025
- Investigating the Robustness of Natural Language Generation from Logical Forms via Counterfactual SamplesChengyuan Liu, Leilei Gan, Kun Kuang, Fei WuEMNLP 2022 · 2 citations
- On the Role of Anticausal Direction in LLM-based Data SynthesisBohan Jiang, Pingchuan Ma, Zhen Tan, Zhuoyu Shi et al.KDD 2026
