Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language Models
Javier González, Aditya V. Nori
Abstract
Recent advances in AI have been significantly driven by the capabilities of large language models (LLMs) to solve complex problems in ways that resemble human thinking. However, there is an ongoing debate about the extent to which LLMs are capable of actual reasoning. Central to this debate are two key probabilistic concepts that are essential for connecting causes to their effects: the probability of necessity (PN) and the probability of sufficiency (PS). This paper introduces a framework that is both theoretical and practical, aimed at assessing how effectively LLMs are able to replicate real-world reasoning mechanisms using these probabilistic measures. By viewing LLMs as abstract machines that process information through a natural language interface, we examine the conditions under which it is possible to compute suitable approximations of PN and PS. Our research marks an important step towards gaining a deeper understanding of when LLMs are capable of reasoning, as illustrated by a series of math examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c0ea598-ce94-4b1e-81c3-297e7f0eba3fCited by top-tier papers8
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang et al.NeurIPS 2025 · 85 citations
- Omitted Variable Bias in Language Models Under Distribution ShiftVictoria Lin, Louis-Philippe Morency, Eli Ben-MichaelICML 2026 · 1 citation
- RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning EvaluationXinnuo Xu, Rachel Lawrence, Kshitij Dubey, Atharva Pandey et al.ICML 2025
- Compositional Causal Reasoning Evaluation in Language ModelsJacqueline R. M. A. Maasch, Alihan Hüyük, Xinnuo Xu, Aditya V. Nori et al.ICML 2025
- Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language ModelsXinyu Yuan, Yan Qiao, Zonghui Wang, Wenzhi CHENICLR 2026
Builds on5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong et al.EMNLP 2023 · 109 citations
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal et al.ACL 2023 · 27 citations
- Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for ReasoningYongqi Tong, Dawei Li, Sizhe Wang, Yujia Wang et al.ACL 2024
Related papers
- Causal Sufficiency and Necessity Improves Chain-of-Thought ReasoningXiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li et al.NeurIPS 2025 · 19 citations
- COLD: Causal reasOning in cLosed Daily activitiesAbhinav Joshi, Areeb Ahmad, Ashutosh ModiNeurIPS 2024 · 11 citations
- A Peek into Token Bias: Large Language Models Are Not Yet Genuine ReasonersBowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang et al.EMNLP 2024 · 27 citations
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological ProfilersYi-Fei Liu, Yi-Long Lu, Di He, Hang ZhangICLR 2026 · 6 citations
- Compositional AI Beyond LLMs: System Implications of Neuro-Symbolic-Probabilistic ArchitecturesZishen Wan, Hanchen Yang, Jiayi Qian, Ritik Raj et al.ASPLOS 2026 · 2 citations
