On the Limit of Language Models as Planning Formalizers
Cassie Huang, Li Zhang
摘要
Large Language Models have been found to create plans that are neither executable nor verifiable in grounded environments. An emerging line of work demonstrates success in using the LLM as a formalizer to generate a formal representation of the planning domain in some language, such as Planning Domain Definition Language (PDDL). This formal representation can be deterministically solved to find a plan. We systematically evaluate this methodology while bridging some major gaps. While previous work only generates a partial PDDL representation, given templated, and therefore unrealistic environment descriptions, we generate the complete representation given descriptions of various naturalness levels. Among an array of observations critical to improve LLMs' formal planning abilities, we note that most large enough models can effectively formalize descriptions as PDDL, outperforming those directly generating plans, while being robust to lexical perturbation. As the descriptions become more natural-sounding, we observe a decrease in performance and provide detailed error analysis. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task PlanningHaoming Ye, Yunxiao Xiao, Cewu Lu, Panpan CaiNeurIPS 2025 · 被引用 9 次
- ESCA: Contextualizing Embodied Agents via Scene-Graph GenerationJiani Huang, Amish Sethi, Matthew Kuo, Mayank Keoliya 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper7
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningLin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 被引用 347 次
- Generalized Planning in PDDL Domains with Pretrained Large Language ModelsTom Silver, Soham Dan, Kavitha Srinivas, Joshua B. Tenenbaum 等AAAI 2024 · 被引用 194 次
- Chain of Thoughtlessness? An Analysis of CoT in PlanningKaya Stechly, Karthik Valmeekam, Subbarao KambhampatiNeurIPS 2024 · 被引用 156 次
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 被引用 123 次
相关 Paper
- Leveraging Environment Interaction for Automated PDDL Translation and Planning with Large Language ModelsSadegh Mahdavi, Raquel Aoki, Keyi Tang, Yanshuai CaoNeurIPS 2024 · 被引用 28 次
- Why Do LLM-based Web Agents Fail? A Hierarchical Planning PerspectiveMohamed Aghzal, Gregory J. Stein, Ziyu YaoACL 2026 · 被引用 5 次
- Language Model as Planner and Formalizer under ConstraintsCassie Huang, Stuti Mohan, Ziyi Yang, Stefanie Tellex 等ACL 2026 · 被引用 3 次
- One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single DemonstrationJinbang Huang, Yixin Xiao, Zhanguang Zhang, Mark Coates 等ICLR 2026 · 被引用 9 次
- SayCanPay: Heuristic Planning with Large Language Models Using Learnable Domain KnowledgeRishi Hazra, Pedro Zuidberg Dos Martires, Luc De RaedtAAAI 2024 · 被引用 74 次
