Open Grounded Planning: Challenges and Benchmark Construction
Shiguang Guo, Ziliang Deng, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun
Abstract
The emergence of large language models (LLMs) has increasingly drawn attention to the use of LLMs for human-like planning. Existing work on LLM-based planning either focuses on leveraging the inherent language generation capabilities of LLMs to produce free-style plans or employs reinforcement learning approaches to learn decision-making for a limited set of actions within restricted environments. However, both approaches exhibit significant discrepancies between the open and executable requirements in real-world planning. In this paper, we propose a new planning task-open grounded planning. The primary objective of open grounded planning is to ask the model to generate an executable plan based on a variable action set, thereby ensuring the executability of the produced plan. To this end, we establish a benchmark for open grounded planning spanning a wide range of domains. Then we test current state-of-the-art LLMs along with five planning approaches, revealing that existing LLMs and methods still struggle to address the challenges posed by grounded planning in open domains. The outcomes of this paper define and establish a foundational dataset for open grounded planning, and shed light on the potential challenges and future directions of LLM-based planning. Our code and datasets are at https://github.com/Shiguang-Guo/ Open-Grounded-Planning * Equal contribution † Corresponding author How to Make Fried Chicken with Tarragon and Buttermilk? Can you check if this address is valid and deliverable? Here's the address: xxx Detect the soft hed boundary of the cake in the image. Restricted Grounded Planning Open Grounded Planning • verifyUSAddress • standardizeUSAddress Closed Action Set Plan verifyUSAddress Open Action Libraries How to Recover from Workout Soreness? Plans Hed Detection On Image Place chicken pieces in a glass or stoneware bowl. Melt Crisco on medium setting. Add chicken to the hot Crisco. Task: How to Activate the Dark Theme on YouTube Method: Using the YouTube App for Android Action Candidate Set: * Close the Tool Options window. * Double click the file. * Do price forecasting. * Click on the blue coloured YOUTUBE STUDIO BETA button. * Open the YouTube app on your iPhone or iPad. * Launch the YouTube app on your Android device. * <other steps>... Steps: 1. Launch the YouTube app on your Android device. 2. Tap on your profile picture. 3. Tap on Settings. 4. Select the General option. 5. Tap on the grey switch, right across Dark theme text. 6. Enjoy YouTube in dark mode <REWRITE PROMPT> INSTRUCTION: You will be given a task, a method to complete the task, a current plan and several candidate actions. Candidate actions are called <Actions in Library>. If no method is specified it will be set to "None". If the current plan is empty, the plan will also be set to "None". Use the actions listed below to refine your current steps to complete your task. Actions marked with <TO BE REPLACED> indicate that the content was not found in the action library, and actions marked with <IN LIB> indicate that they are in the action library. You need to analyze which actions in the provided action library can be added to the action list and replace some or all of the actions marked with <TO BE REPLACED>. We encourage you to add more <TO BE REPLACED> content to complete these steps. You can do the following: * Replace any number of <TO BE REPLACED>-like operations with any number of <IN LIB> operations. * Replace any number of <IN LIB> operations with any number of <IN LIB> operations as the latter are better suited to the task. * Replace any number of <IN LIB> operations with more <IN LIB> operations. * Insert any number of <IN LIB> operations that differ from existing steps. * Insert any number of <TO BE REPLACED> operations to fill in missing content between steps. * Remove any number of redundant <IN LIB> operations. * Remove any number of redundant <TO BE REPLACED> operations. * Remove any number of overly verbose <IN LIB> operations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PlanGenLLMs: A Modern Survey of LLM Planning CapabilitiesHui Wei, Zihao Zhang, Shenghua He, Tian Xia et al.ACL 2025 · 78 citations
- Structured Multi-step Jailbreaking under a Hamiltonian Generative FormulationZihan Zhou, Yang Zhou, Jianghai Yu, Lingjuan Lyu et al.ICML 2026
Builds on11
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
Related papers
- Symbolic Planning and Code Generation for Grounded DialogueJustin T. Chiu, Wenting Zhao, Derek Chen, Saujas Vaduguru et al.EMNLP 2023
- On the Limit of Language Models as Planning FormalizersCassie Huang, Li ZhangACL 2025
- Visual Programming for Zero-Shot Open-Vocabulary 3D Visual GroundingZhihao Yuan, Jinke Ren, Chun-Mei Feng, Hengshuang Zhao et al.CVPR 2024 · 19 citations
- Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill LearningShaohui Peng, Xing Hu, Qi Yi, Rui Zhang et al.AAAI 2024 · 4 citations
- On Grounded Planning for Embodied Tasks with Language ModelsBill Yuchen Lin, Chengsong Huang, Qian Liu, Wenda Gu et al.AAAI 2023 · 52 citations
