PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool Learning
Zihan Zheng, Tianle Cui, Chuwen Xie, Jiahui Pan, Qianglong Chen, Lewei He
Abstract
One of the research focuses of large language models (LLMs) is the ability to generate action plans. Recent studies have revealed that the performance of LLMs can be significantly improved by integrating external tools. Based on this, we propose a benchmark framework called PlanningArena, which aims to simulate real application scenarios and provide a series of apps and API tools that may be involved in the actual planning process. This framework adopts a modular task structure and combines user portrait analysis to evaluate the ability of LLMs in correctly selecting tools, logical reasoning in complex scenarios, and parsing user information. In addition, we deeply diagnose the task execution effect of LLMs from both macro and micro levels. The experimental results show that even the most outstanding GPT-4o and DeepSeekV3 models only achieved a total score of 56.5% and 41.9% in PlanningArena, respectively, indicating that current LLMs still face challenges in logical reasoning, context memory, and tool calling when dealing with different structures, scenarios, and their complexity. Through this benchmark, we further explore the path to optimize LLMs to perform planning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f6d3ee7-ecde-494b-a7c7-38d206e4f0bcCited by top-tier papers2
- NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI TasksZihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang et al.ACL 2026
- BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering ConstructionTian Xia, Tianrun Gao, Wenhao Deng, Long Wei et al.ICML 2026
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu et al.ICML 2024 · 376 citations
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningLin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 347 citations
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu et al.ACL 2023 · 249 citations
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong et al.EMNLP 2023 · 109 citations
Related papers
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia et al.EMNLP 2024 · 2 citations
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song et al.EMNLP 2023 · 72 citations
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding TasksHanwen Xu, Xuyao Huang, Yuzhe Liu, Zhijie DengACL 2026 · 2 citations
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain ScenariosAelim Ahn, Sooyeon Lee, Hyosun Wang, Chiwan Park et al.ICLR 2026
