CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans
Yash Kumar Lal, Vanya Cohen, Nathanael Chambers, Niranjan Balasubramanian, Raymond J. Mooney
Abstract
Understanding the abilities of LLMs to reason about natural language plans, such as instructional text and recipes, is critical to reliably using them in decision-making systems.A fundamental aspect of plans is the temporal order in which their steps need to be executed, which reflects the underlying causal dependencies between them.We introduce CAT-BENCH, a benchmark of Step Order Prediction questions, which test whether a step must necessarily occur before or after another in cooking recipe plans.We use this to evaluate how well frontier LLMs understand causal and temporal dependencies.We find that SOTA LLMs are underwhelming (best zero-shot is only 0.59 in F1), and are biased towards predicting dependence more often, perhaps relying on temporal order of steps as a heuristic.While prompting for explanations and using few-shot examples improve performance, the best F1 result is only 0.73.Further, human evaluation of explanations along with answer correctness show that, on average, humans do not agree with model reasoning.Surprisingly, we also find that explaining after answering leads to better performance than normal chain-of-thought prompting, and LLM answers are not consistent across questions about the same step pairs.Overall, results show that LLMs' ability to detect dependence between steps has significant room for improvement. * Equal ContributionAlmond Flour Chocolate Cake Step 6: Stir in ground almonds.Step 7: Add half flour and half milk.Step 8: Use wooden spoon to stir. Step 12: Whip cream till stiff peaks Q: Must Step 6 happen before Step 8? Questions about dependent steps Q: Must Step 7 happen after Step 6? Questions about non-dependent steps A: Yes, all ingredients have to be in bowl before
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMsKarin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu et al.ACL 2025
- Benchmarking Agentic Workflow GenerationShuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang et al.ICLR 2025
- How2Everything: Mining the Web for How-to Procedures to Evaluate and Improve LLMsYapei Chang, Kyle Lo, Mohit Iyyer, Luca SoldainiICML 2026
- Transparent and Coherent Procedural Mistake DetectionShane Storks, Itamar Bar-Yossef, Yayuan Li, Zheyuan Zhang et al.EMNLP 2025
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- The Generative AI Paradox: "What It Can Create, It May Not Understand"Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman et al.ICLR 2024 · 116 citations
- A Dataset for Tracking Entities in Open Domain Procedural TextNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal et al.EMNLP 2020 · 38 citations
- Analogous Process Structure Induction for Sub-event Sequence PredictionHongming Zhang, Muhao Chen, Haoyu Wang, Yangqiu Song et al.EMNLP 2020 · 31 citations
Related papers
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha et al.ICML 2026
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language ModelsHyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok LimEMNLP 2025
- CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in VideosXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi Huang et al.AAAI 2026 · 7 citations
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
