Language Model as Planner and Formalizer under Constraints
Cassie Huang, Stuti Mohan, Ziyi Yang, Stefanie Tellex, Li Zhang
摘要
LLMs have been widely used in planning, either as planners to generate action sequences end-to-end, or as formalizers to represent the planning domain and problem in a formal language that can derive plans deterministically. However, both lines of work rely on standard benchmarks that include only generic and simplistic environmental specifications, leading to potential overestimation of the planning ability of LLMs and safety concerns in downstream tasks. We bridge this gap by augmenting widely used planning benchmarks with manually annotated, fine-grained, and rich natural language constraints spanning four formally defined categories. Over 4 state-of-the-art reasoning LLMs, 4 formal languages, and 4 datasets, we show that the introduction of one-sentence constraints consistently halves performance, indicating current LLMs' lack of robustness and an avenue for future research. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksBill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman 等NeurIPS 2023 · 被引用 244 次
- Chain of Thoughtlessness? An Analysis of CoT in PlanningKaya Stechly, Karthik Valmeekam, Subbarao KambhampatiNeurIPS 2024 · 被引用 156 次
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 被引用 123 次
- PlanGenLLMs: A Modern Survey of LLM Planning CapabilitiesHui Wei, Zihao Zhang, Shenghua He, Tian Xia 等ACL 2025 · 被引用 78 次
- LLM+AL: Bridging Large Language Models and Action Languages for Complex Reasoning About ActionsAdam Ishay, Joohyung LeeAAAI 2025 · 被引用 12 次
相关 Paper
- On the Limit of Language Models as Planning FormalizersCassie Huang, Li ZhangACL 2025
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 被引用 35 次
- Planning in the Dark: LLM-Symbolic Planning Pipeline Without ExpertsSukai Huang, Nir Lipovetzky, Trevor CohnAAAI 2025 · 被引用 19 次
- LLMs Can Plan Only If We Tell ThemBilgehan Sel, Ruoxi Jia, Ming JinICLR 2025
- Logical forms complement probability in understanding language model (and human) performanceYixuan Wang, Freda ShiACL 2025 · 被引用 2 次
