CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans
Yash Kumar Lal, Vanya Cohen, Nathanael Chambers, Niranjan Balasubramanian, Raymond J. Mooney
摘要
Understanding the abilities of LLMs to reason about natural language plans, such as instructional text and recipes, is critical to reliably using them in decision-making systems.A fundamental aspect of plans is the temporal order in which their steps need to be executed, which reflects the underlying causal dependencies between them.We introduce CAT-BENCH, a benchmark of Step Order Prediction questions, which test whether a step must necessarily occur before or after another in cooking recipe plans.We use this to evaluate how well frontier LLMs understand causal and temporal dependencies.We find that SOTA LLMs are underwhelming (best zero-shot is only 0.59 in F1), and are biased towards predicting dependence more often, perhaps relying on temporal order of steps as a heuristic.While prompting for explanations and using few-shot examples improve performance, the best F1 result is only 0.73.Further, human evaluation of explanations along with answer correctness show that, on average, humans do not agree with model reasoning.Surprisingly, we also find that explaining after answering leads to better performance than normal chain-of-thought prompting, and LLM answers are not consistent across questions about the same step pairs.Overall, results show that LLMs' ability to detect dependence between steps has significant room for improvement. * Equal ContributionAlmond Flour Chocolate Cake Step 6: Stir in ground almonds.Step 7: Add half flour and half milk.Step 8: Use wooden spoon to stir. Step 12: Whip cream till stiff peaks Q: Must Step 6 happen before Step 8? Questions about dependent steps Q: Must Step 7 happen after Step 6? Questions about non-dependent steps A: Yes, all ingredients have to be in bowl before
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMsKarin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu 等ACL 2025
- Benchmarking Agentic Workflow GenerationShuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang 等ICLR 2025
- How2Everything: Mining the Web for How-to Procedures to Evaluate and Improve LLMsYapei Chang, Kyle Lo, Mohit Iyyer, Luca SoldainiICML 2026
- Transparent and Coherent Procedural Mistake DetectionShane Storks, Itamar Bar-Yossef, Yayuan Li, Zheyuan Zhang 等EMNLP 2025
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- The Generative AI Paradox: "What It Can Create, It May Not Understand"Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman 等ICLR 2024 · 被引用 116 次
- A Dataset for Tracking Entities in Open Domain Procedural TextNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal 等EMNLP 2020 · 被引用 38 次
- Analogous Process Structure Induction for Sub-event Sequence PredictionHongming Zhang, Muhao Chen, Haoyu Wang, Yangqiu Song 等EMNLP 2020 · 被引用 31 次
相关 Paper
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha 等ICML 2026
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language ModelsHyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok LimEMNLP 2025
- CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in VideosXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi Huang 等AAAI 2026 · 被引用 7 次
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura 等ICLR 2025
