Exploring the Capacity of Pretrained Language Models for Reasoning about Actions and Change
Weinan He, Canming Huang, Zhanhao Xiao, Yongmei Liu
Abstract
Reasoning about actions and change (RAC) is essential to understand and interact with the ever-changing environment. Previous AI research has shown the importance of fundamental and indispensable knowledge of actions, i.e., preconditions and effects. However, traditional methods rely on logical formalization which hinders practical applications. With recent transformer-based language models (LMs), reasoning over text is desirable and seemingly feasible, leading to the question of whether LMs can effectively and efficiently learn to solve RAC problems. We propose four essential RAC tasks as a comprehensive textual benchmark and generate problems in a way that minimizes the influence of other linguistic requirements (e.g., grounding) to focus on RAC. The resulting benchmark, TRAC, encompassing problems of various complexities, facilitates a more granular evaluation of LMs, precisely targeting the structural generalization ability much needed for RAC. Experiments with three high-performing transformers indicate that additional efforts are needed to tackle challenges raised by TRAC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cae8fcc-9135-400f-a7e6-7bd0455aecc8Cited by top-tier papers4
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- ACPBench Hard: Unrestrained Reasoning about Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiICLR 2026 · 9 citations
- ActionReasoningBench: Reasoning about Actions with and without Ramification ConstraintsDivij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son et al.ICLR 2025
- MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation DatasetWeiqi Wang, Yangqiu SongACL 2025
Builds on3
- PRover: Proof Generation for Interpretable Reasoning over RulesSwarnadeep Saha, Sayan Ghosh, Shashank Srivastava, Mohit BansalEMNLP 2020 · 3 citations
- Implicit Representations of Meaning in Neural Language ModelsBelinda Z. Li, Maxwell I. Nye, Jacob AndreasACL 2021
- PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D WorldRowan Zellers, Ari Holtzman, Matthew E. Peters, Roozbeh Mottaghi et al.ACL 2021
Related papers
- ScienceWorld: Is your Agent Smarter than a 5th Grader?Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj AmmanabroluEMNLP 2022 · 1 citation
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist ModelsFangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang et al.ICML 2026
- Evaluating Relational Reasoning in LLMs with RELLukas Fesser, Yasha Ektefaie, Ada Fang, Sham Kakade et al.ICML 2026
