COLD: Causal reasOning in cLosed Daily activities
Abhinav Joshi, Areeb Ahmad, Ashutosh Modi
Abstract
Large Language Models (LLMs) have shown state-of-the-art performance in a variety of tasks, including arithmetic and reasoning; however, to gauge the intellectual capabilities of LLMs, causal reasoning has become a reliable proxy for validating a general understanding of the mechanics and intricacies of the world similar to humans. Previous works in natural language processing (NLP) have either focused on open-ended causal reasoning via causal commonsense reasoning (CCR) or framed a symbolic representation-based question answering for theoretically backed-up analysis via a causal inference engine. The former adds an advantage of real-world grounding but lacks theoretically backed-up analysis/validation, whereas the latter is far from real-world grounding. In this work, we bridge this gap by proposing the COLD (Causal reasOning in cLosed Daily activities) framework, which is built upon human understanding of daily real-world activities to reason about the causal nature of events. We show that the proposed framework facilitates the creation of enormous causal queries ( 9 million) and comes close to the mini-turing test, simulating causal reasoning to evaluate the understanding of a daily real-world task. We evaluate multiple LLMs on the created causal queries and find that causal reasoning is challenging even for activities trivial to humans. We further explore (the causal reasoning abilities of LLMs) using the backdoor criterion to determine the causal strength between events.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ed077bd-af7d-4721-ae58-dcdb2f86182eCited by top-tier papers5
- Geometry of Decision Making in Language ModelsAbhinav Joshi, Divyanshu Bhatt, Ashutosh ModiNeurIPS 2025 · 12 citations
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 9 citations
- Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-AuditingWenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu et al.ACL 2026 · 3 citations
- Uncertainty in Causality: A New FrontierShaobo Cui, Luca Mouchel, Boi FaltingsACL 2025
- LLMs Struggle to Balance Reasoning and World Knowledge in Causal Narrative UnderstandingKhurram Yamin, Shantanu Gupta, Gaurav R. Ghosal, Zachary C. Lipton et al.ICLR 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas et al.ICLR 2023 · 60 citations
- Leveraging Large Language Models for Multiple Choice Question AnsweringJoshua Robinson, David WingateICLR 2023 · 40 citations
Related papers
- CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language ModelsZhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele et al.NeurIPS 2023 · 74 citations
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?Haoang Chi, He Li, Wenjing Yang, Feng Liu et al.NeurIPS 2024 · 124 citations
- Language Agents Meet Causality - Bridging LLMs and Causal World ModelsJohn Gkountouras, Matthias Lindemann, Phillip Lippe, Efstratios Gavves et al.ICLR 2025
- Com² : A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language ModelsKai Xiong, Xiao Ding, Yixin Cao, Yuxiong Yan et al.ACL 2025
- Mitigating Hallucinations in Large Language Models via Causal ReasoningYuangang Li, Yiqing Shen, Yi Nian, Jiechao Gao et al.AAAI 2026 · 1 citation
