Exploring the Capacity of Pretrained Language Models for Reasoning about Actions and Change
Weinan He, Canming Huang, Zhanhao Xiao, Yongmei Liu
摘要
Reasoning about actions and change (RAC) is essential to understand and interact with the ever-changing environment. Previous AI research has shown the importance of fundamental and indispensable knowledge of actions, i.e., preconditions and effects. However, traditional methods rely on logical formalization which hinders practical applications. With recent transformer-based language models (LMs), reasoning over text is desirable and seemingly feasible, leading to the question of whether LMs can effectively and efficiently learn to solve RAC problems. We propose four essential RAC tasks as a comprehensive textual benchmark and generate problems in a way that minimizes the influence of other linguistic requirements (e.g., grounding) to focus on RAC. The resulting benchmark, TRAC, encompassing problems of various complexities, facilitates a more granular evaluation of LMs, precisely targeting the structural generalization ability much needed for RAC. Experiments with three high-performing transformers indicate that additional efforts are needed to tackle challenges raised by TRAC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 被引用 35 次
- ACPBench Hard: Unrestrained Reasoning about Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiICLR 2026 · 被引用 9 次
- ActionReasoningBench: Reasoning about Actions with and without Ramification ConstraintsDivij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son 等ICLR 2025
- MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation DatasetWeiqi Wang, Yangqiu SongACL 2025
它引用的顶会 Paper3
- PRover: Proof Generation for Interpretable Reasoning over RulesSwarnadeep Saha, Sayan Ghosh, Shashank Srivastava, Mohit BansalEMNLP 2020 · 被引用 3 次
- Implicit Representations of Meaning in Neural Language ModelsBelinda Z. Li, Maxwell I. Nye, Jacob AndreasACL 2021
- PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D WorldRowan Zellers, Ari Holtzman, Matthew E. Peters, Roozbeh Mottaghi 等ACL 2021
相关 Paper
- ScienceWorld: Is your Agent Smarter than a 5th Grader?Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj AmmanabroluEMNLP 2022 · 被引用 1 次
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu 等ICML 2026
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist ModelsFangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang 等ICML 2026
- Evaluating Relational Reasoning in LLMs with RELLukas Fesser, Yasha Ektefaie, Ada Fang, Sham Kakade 等ICML 2026
