ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning
Harsha Kokel, Michael Katz, Kavitha Srinivas, Shirin Sohrabi
摘要
We introduce ACPBench Hard, a dataset of generative, open-ended questions which LLM models needs to answer in order to plan. Models that perform well on these tasks could in principle be integrated into a planner or be used directly as a policy. We discuss the complexity of these tasks as well as the complexity of validating the correctness of their answers and present validation algorithms for each task. Equipped with these validators, we test the performance of a variety of models on our tasks and find that for most of these tasks, the performance of even the largest models is still subpar. The models do not possess even the most basic capability of identifying which actions can be performed in a given state. No model outperforms any other on our proposed tasks and, with a few exceptions, all tested language models score below 65%, indicating that even the current frontier language models as well as so-called reasoning models have a long way to go before they can reliably reason about planning.
ACPBench Hard collection is publicly available, see https://ibm.github.io/ACPBench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 被引用 509 次
- Reshaping Diverse PlanningMichael Katz, Shirin SohrabiAAAI 2020 · 被引用 55 次
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 被引用 38 次
- Top-Quality Planning: Finding Practically Useful Sets of Best PlansMichael Katz, Shirin Sohrabi, Octavian UdreaAAAI 2020 · 被引用 36 次
相关 Paper
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 被引用 35 次
- BIG-Bench Extra HardMehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch 等ACL 2025
- TCP: a Benchmark for Temporal Constraint-Based PlanningZifeng Ding, Sikuan Yan, Moy Yuan, Xianglong Hu 等EMNLP 2025
- Logical forms complement probability in understanding language model (and human) performanceYixuan Wang, Freda ShiACL 2025 · 被引用 2 次
- NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity ClassesLizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling 等ACL 2024 · 被引用 8 次
